Chapter 1 - Introduction
What is this book about?
This series of posts is about using Rust for data engineering tasks for people who are already familiar with Python and are curious about Rust. It will not cover every aspect of Rust, or of Python. Instead, it aims to give practical examples of how common engineering tasks done in Python might be done in Rust, along with representative benchmarks.
We will cover a variety of topics, including:
- Getting data from an API
- Scraping a website
- Parsing data and using structs
- Data Transformation
- Writing Data
- Other topics as I feel like it
This book is not an introduction to either Rust or Python. There are many great resources to both out there. If you are not familiar with Python, the official Python Tutorial is a great starting point.
Should I use Rust for Data Engineering?
Probably not. Rust is a great language, it is fun, it is pleasant to use, and it is fast. But choosing a language for a project is more than choosing a language that is fun. There are cautionary tales about using Rust at a startup, and I think they are worth reading.
There are many reasons why you might not want to use Rust for data engineering. The first is that Rust is not as mature as Python. There are many libraries that are missing. For example, as of this writing, there are no Snowflake libraries for querying data in that warehouse. Most people do not know Rust, and it is harder to hire and harder to train people.
There may be good reasons to use Rust for data engineering however. When it comes to cost and performance, Rust is clearly faster than Python for many types of tasks. Memory usage is also much lower, which can be important when you are constrained by small IoT devices, for example.
I can’t tell you when to use Rust and when to use Python, but I do believe that by understanding both languages, their merits and pitfalls, you will be better positioned to make that decision for yourself.
Why Should I Learn Rust?
Because it is fun to learn new things. I can’t promise you that anything you learn here will ever have a material impact on your life or career. But if you enjoy learning and tinkering, then you might want to tinker with this. If you are like me, and you like learning for learning’s sake, then you will enjoy this experience too. I learned vim and lua not because it was useful, but because I was curious about it. I did end up benefiting from it, but I never approached it from a purely utilitarian perspective. There are better ways to spend your time if your goal is purely career advancement.
But, if you are curious about Rust, and if you like to have fun, then I think you will be pleasantly surprised by what Rust has to offer.
Prerequisites
Installing Rust and Python
You will need Rust and Python installed to follow along with the examples here.
For Rust, go to rustup.rs. This gives you rustc and
cargo, which is all we need.
For Python, I recommend uv. It replaces the pile
of tools I used to reach for — pyenv for interpreters, virtualenv for
environments, pip for packages — with a single one, and it will install the
right Python for you if you don’t have it.
# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh
A typical project looks like this:
# Create a project
uv init somepyproj
cd somepyproj
# Pin the Python version for this project
uv python pin 3.13
# Add a dependency
uv add polars
# Run something inside the project environment
uv run python -m somepyproj.main
uv creates the virtual environment for you the first time you need one, so
there is no separate “activate the venv” step. uv add writes the dependency
into pyproject.toml and records the exact resolved version in uv.lock,
which is the Python equivalent of Rust’s Cargo.lock.
A note on versions
Both sample projects in this book pin their dependencies exactly and commit their lockfiles, so the code you build should behave the same as the code the benchmarks were run against. At the time of writing that means:
| Version | |
|---|---|
| Rust | 1.97.0 |
| Python | 3.13 |
| polars (Rust) | 0.54.4 |
| polars (Python) | 1.43.0 |
| pandas | 3.0.5 |
Both of these ecosystems move quickly, and polars in particular has changed
its API substantially over the years. If you are reading this well after it was
written, expect some drift.
Installing the Code
The code for this book can be found here: https://github.com/PedramNavid/rust-for-data
git clone git@github.com:PedramNavid/rust-for-data.git
The Rust examples live in wxrs, and the Python examples live in wxpy.
Fetching from an API
One of the simplest examples to start with is fetching data from an API endpoint. This is often the beginning of many data pipeline journeys.
In our first case, we will use the OpenWeatherMap API to fetch the current weather in a configurable location by providing a latitude and longitude on the command line.
You will need to sign up for a free account to get an API key, once you’ve signed up, create an API key.
Starting a Project
One of the first differences between Rust and Python you will experience is through initializing a project.
Rust
In Rust, this is as simple as running
# Create the project
cargo init wxrs
# Add a dependency
cd wxrs
cargo add reqwest --features blocking
This will create a new directory called wxrs with a Hello World example.
It will also add the reqwest crate to our dependencies, similar to pip install.
Unlike a bare pip install though, this will also update Cargo.toml with our
dependency, and create a Cargo.lock file that pins the reqwest crate to a
specific version. uv add does the same thing on the Python side.
The --features flag is used to express optional compilation features. Reqwest
has several options, described in the crate’s documentation.
We will use the blocking feature, which will allow us to use the blocking API.
It gives us a simpler interface to reqwest instead of futures that require
an async runtime. We will eventually use async to show the power of Rust’s
fearless concurrency.
We are not using any specific features, but Python too allows optional features,
for example pip install snowflake-connector-python[pandas].
# Cargo.toml
[package]
name = "wxrs"
version = "0.1.0"
edition = "2021"
# Dependency versions are pinned exactly and Cargo.lock is committed, so the
# code and benchmarks printed in the book match what you get when you build it.
# See more keys and their definitions at https://doc.rust-lang.org/cargo/reference/manifest.html
[dependencies]
polars = { version = "=0.54.4", features = ["lazy", "csv", "strings"] }
reqwest = { version = "=0.13.4", features = ["blocking", "json"] }
serde = { version = "=1.0.229", features = ["derive"] }
serde_json = "=1.0.151"
Python
This used to be the part where I explained that Python makes you do all of this
by hand: create the directory, create a virtual environment, hand-write a
pyproject.toml, name your dependencies, then install the package locally. It
was a genuinely unflattering comparison.
uv has since closed most of that gap, and the Python side now looks a lot
like the Rust side:
# Create the project
uv init wxpy
cd wxpy
# Add a dependency
uv add requests
Like cargo add, uv add writes the dependency into pyproject.toml and
records the exact resolved version in a lockfile — uv.lock here, Cargo.lock
there.
# pyproject.toml
[project]
name = "wxpy"
version = "0.0.1"
requires-python = ">=3.11"
# Versions are pinned exactly, and uv.lock is committed, so the code and
# benchmarks printed in the book match what you get when you run it yourself.
dependencies = [
"requests==2.34.2",
"polars==1.43.0",
"pandas==3.0.5",
]
Now, admittedly we can skip all of the above steps, create a random file
anywhere we want and run it with python myfile.py, but the goal here is to
build a more stable distribution that can be packaged, shared, and tested.
It is worth being honest about what this does to the comparison. For years the Rust story here was simply better, and that was a real argument in Rust’s favour. It isn’t much of one anymore.
Fetching the Weather
Now that we have a project, let’s fetch the weather. To fetch from an API,
we will use the requests package in Python and the reqwest package in Rust.
We will read the API key from the environment variable, and get the latitude and longitude from the command line arguments.
Given that I’m in California, it only makes sense to start with the Air Pollution API.
Python
In Python, we’ll create a folder for this chapter to keep code organized, and then run that file directly.
mkdir wxpy/wxpy/ch3
# wxpy/wxpy/ch3/fetch_api.py
import os
import sys
import requests
API_KEY = os.getenv("OWM_APPID")
def get_air_pollution(lat, lon):
url = f"http://api.openweathermap.org/data/2.5/air_pollution?lat={lat}&lon={lon}&appid={API_KEY}"
body = requests.get(url).text
return body
if __name__ == "__main__":
usage = f"Usage: python {__file__} <lat> <lon>"
if not API_KEY:
print("Please set OWM_APPID environment variable")
sys.exit(1)
if len(sys.argv) != 3:
print(usage)
sys.exit(1)
lat = sys.argv[1]
lon = sys.argv[2]
body = get_air_pollution(lat, lon)
print(body)
Rust
In Rust, usually we’ll have a main.rs file that runs our code, with
additional code imported as modules from other files. There’s a great
convention for package layouts in Rust.
But since we want to execute our code directly, and we’ll have multiple
binaries, we’ll create a bin folder and save the code for ch3 there.
mkdir wxrs/src/bin
// wxrs/src/bin/ch3.rs
pub fn get_air_pollution(lat: f32, lon: f32) -> String {
let api_key = std::env::var("OWM_APPID").expect(
"Environment Variable OWM_APPID not set. Please set it to your
OpenWeatherMap API key. https://home.openweathermap.org/api_keys",
);
let url = format!(
"http://api.openweathermap.org/data/2.5/air_pollution?lat={}&lon={}&appid={}",
lat, lon, api_key
);
reqwest::blocking::get(url)
.expect("request failed")
.text()
.expect("body failed")
}
pub fn main() {
let usage = format!("Usage: {} [lat] [lon]", std::env::args().next().unwrap());
let lat = std::env::args()
.nth(1)
.expect(&usage)
.parse::<f32>()
.expect(&usage);
let lon = std::env::args()
.nth(2)
.expect(&usage)
.parse::<f32>()
.expect(&usage);
let body = get_air_pollution(lat, lon);
println!("{}", body);
}
Running the Program
Running the program is simple in both languages. We’ll provide the latitude and longitude of beautiful Fairfax, CA, birthplace of mountain biking, and nestled in the foothills of Mount Tamalpais.
Google gives the coordinates as 37.9871 and -122.5889
Python
In Python, we can use -m to run the module directly. uv run takes care of
creating and using the project environment.
# in wxpy
# export OWM_APPID=your-api-key
uv run python -m wxpy.ch3.fetch_api 37.9871 -122.5889
> {"coord":{"lon":-122.5889,"lat":37.9871},"list":[{"main":{"aqi":2},"components":{"co":178.58,"no":0.1,"no2":0.47,"o3":70.81,"so2":0.64,"pm2_5":2.58,"pm10":4.18,"nh3":0},"dt":1687221287}]}
Rust
In Rust, we must first compile the program before running it. If we run
cargo build Rust will create a binary for us in ./target/debug/wxrs
We can also compile and run it with one command cargo run
When using cargo build Rust will build a debug version of our application in
./target/debug for both the main.rs file which will be named wxrs as
well for any files located in src/bin, such as ch3.rs
# in wxrs/
cargo build
./target/debug/ch3 37.9871 -122.5889
> {"coord":{"lon":-122.5889,"lat":37.9871},"list":[{"main":{"aqi":2},"components":{"co":178.58,"no":0.1,"no2":0.47,"o3":70.81,"so2":0.64,"pm2_5":2.58,"pm10":4.18,"nh3":0},"dt":1687221453}]}
# or
cargo run --bin ch3 37.9871 -122.5889
> {"coord":{"lon":-122.5889,"lat":37.9871},"list":[{"main":{"aqi":2},"components":{"co":178.58,"no":0.1,"no2":0.47,"o3":70.81,"so2":0.64,"pm2_5":2.58,"pm10":4.18,"nh3":0},"dt":1687221453}]}
Discussion
Looking at both programs, we can see a fairly similar approach to solving this problem.
Both programs use an external library or crate (not-so-coincidentally named request/reqwest).
In both programs, we’ve created a function that takes a latitude and longitude, fetches the results from an API and returns the results as text. We’ll cover handling structured data from JSON soon.
Types
One obvious difference is that in Rust, we declare the types of the lat and lon arguments, and in Python we do not. The trouble with talking about types is that it inevitably leads to a discussion of memory, which can devolve into a conversation around null pointer references, which we will largely avoid until the next chapter, but here’s a light introduction.
In the Rust code, we’ve very explicitly defined the types for our function:
#![allow(unused)]
fn main() {
pub fn get_air_pollution(lat: f32, lon: f32) -> String {
}
Both lat and lon are f32 or 32-bit floats. These are floating-point numbers
that take exactly 32-bits of memory. The compiler knows exactly how much space
to reserve for these values: 32-bits, or 4-bytes.
Given that lat and lon doesn’t require much precision beyond a few
decimals, f32 seems like the best choice for our code. We could even opt for
greater precision by using a 64-bit float or f64 in Rust which would take 8
bytes of memory.
Because we know exactly how much memory we need for these variables, and we know it at compile time, Rust is able to store these values on the stack. Stack allocation is cheap: it’s a pointer bump, and the memory is reclaimed automatically when the function returns.
In Python, we don’t know what these arguments will be until runtime, because Python will accept anything in this function.
def get_air_pollution(lat, lon):
url = f"http://api.openweathermap.org/data/2.5/air_pollution?lat={lat}&lon={lon}&appid={API_KEY}"
We could pass it a string, numbers, another function, or even None.
>>> def join_two(a, b):
... return f"a+b={a}+{b}"
...
>>> join_two(1,2)
'a+b=1+2'
>>> join_two(None, None)
'a+b=None+None'
>>> join_two(join_two, join_two)
'a+b=<function join_two at 0x7f8f7f7de980>+<function join_two at 0x7f8f7f7de980>'
>>> join_two(join_two, join_two(join_two, join_two))
'a+b=<function join_two at 0x7f8f7f7de980>+a+b=<function join_two at 0x7f8f7f7de980>+<function join_
two at 0x7f8f7f7de980>'
Even the url line will not fail, because in Python duck-typing allows
us great flexibility in what we do with variables. We can pass numbers into an
f-string for concatenation just as easily as we can pass characters.
Python will allocate these values on the heap, and it turns out that Python allocates about 24 bytes for each float there. The actual values are stored in a private heap.
Now, the difference between Python’s 24 bytes and Rust’s 4 is trivial for an application such as this, and even on the most memory-constrained devices it’s not worth noting. But it’s important to know that heap allocation is slower, and even small applications may iterate over millions of values. Small differences can add up.
You might then ask yourself: what about mypy? Doesn’t that give us typing? Mypy is a static type checker, but it doesn’t change the underlying compilation of Python code. It can provide hints as to what you expect the types to be, but it doesn’t change how memory is allocated.
Handling Errors
Another subtle but important difference is the handling of errors. In Python, errors are handled as exceptions that are caught. Knowing when to catch an exception is mostly an art. It’s difficult to know which functions throw exceptions, what exceptions to expect, and when to deal with them.
In Rust, errors are handled as values that are returned. This is a much more
explicit approach, and it’s easier to know what errors to expect and how to
handle them. A function’s signature tells you up front that it can fail, and
Result is marked #[must_use], so ignoring one is a compiler warning rather
than something you silently overlook. This is a huge benefit to Rust, and it’s
one of the reasons why Rust is so reliable.
Let’s take a closer look at what happens with bad input. If we run the Python program with arguments that aren’t numbers at all, nothing complains locally:
uv run python -m wxpy.ch3.fetch_api nice birds
> {"cod":"400","message":"wrong latitude"}
./target/debug/ch3 nice birds
> thread 'main' (10541501) panicked at src/bin/ch3.rs:24:10:
Usage: ./target/debug/ch3 [lat] [lon]: ParseFloatError { kind: Invalid }
note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace
What happened here?
In Python, we didn’t check that our input is valid, and so the application sent incorrect values to the API, which returned an error message. Fortunately, we get a million API requests for free a month, so this one doesn’t cost us much.
In Rust, the application panicked because it could not parse the inputs we provided
as a float. On lines 20-24 we call parse on the arguments, and we expect them
to be floats.
#![allow(unused)]
fn main() {
let lat = std::env::args()
.nth(1)
.expect(&usage)
.parse::<f32>()
.expect(&usage);
}
The expect method tells Rust that if parse failed to convert the input,
then the application must panic. We output the usage message and exit.
You will see expect and its cousin unwrap used frequently in Rust. They are
useful for debugging, but they are not the best way to handle errors. We’ll
cover error handling in more detail soon.
Benchmarks
Let me preface this by saying speed isn’t everything. No doubt someone familiar in Python will spend far more time learning Rust than they might ever save by running a slightly more optimized program. But it is nice to get a sense of the difference, and to watch how it changes as the programs get less trivial.
Let’s use hyperfine to benchmark the two programs. We’ll run each program
10 times and take the average. Before we benchmark the Rust application,
we’ll compile it using --release which builds a release rather than a
debug version and it should provide us with a faster application.
cargo build --release
Every benchmark in this book is generated by the Makefile in benchmarks/,
so you can reproduce them yourself:
# from the repo root
make build-release
OWM_APPID=your-api-key make benchmarks
Under the hood that is just hyperfine comparing the two binaries:
hyperfine --warmup 5 \
'../wxrs/target/release/ch3 30 -140' \
'../wxpy/.venv/bin/python ../wxpy/wxpy/ch3/fetch_api.py 30 -140' \
--export-markdown ch3_fetch_api.md
| Command | Mean [ms] | Min [ms] | Max [ms] | Relative |
|---|---|---|---|---|
../wxrs/target/release/ch3 30 -140 | 108.6 ± 9.0 | 96.0 | 132.0 | 1.00 |
../wxpy/.venv/bin/python ../wxpy/wxpy/ch3/fetch_api.py 30 -140 | 141.2 ± 10.0 | 133.3 | 181.4 | 1.30 ± 0.14 |
On my system, the Python application took an average of 141ms to complete, while the Rust application was 1.3x faster at 109ms. Memory consumption was also lower in Rust, with the Python application using 33MB against only 10MB in Rust.
That 1.3x is worth picking apart, because it is not really telling us anything
about how fast either language fetches a URL. Both programs spend most of their
time waiting on the same network. The interesting column is the CPU time
hyperfine reports: about 10ms of user time for Rust against about 46ms for
Python. That ~32ms difference is almost exactly the wall-clock gap between the
two, and it is mostly the cost of starting a Python interpreter and importing
requests.
In other words, for a program this small, we are benchmarking startup. That is a real cost if you are invoking a script thousands of times from a scheduler, and completely irrelevant if you are running one long-lived process.
Again, this is a trivial application with trivial requirements and performance is not a key factor in deciding what language to build. But as we build more intensive applications we’ll keep an eye on memory and performance to see how the gap changes.
Summary
In this chapter we’ve built a simple application that fetches data from an API and returns the results. We’ve seen how Rust and Python differ in their approach to handling errors and types, and we’ve seen how Rust can be faster and more memory efficient than Python.
Serializing Data
In the last chapter we fetched data from the OpenWeather API in order to get Air Pollution data. The astute observer will have noticed that we parsed the response as pure text, although the response was in JSON format.
The goal of this chapter is to walk through how we would take raw data and serialize it into a structured data format, such as JSON.
We’ll dive into theory in a little but let’s start with practice.
Serialization
Serialization is the process of taking data and encoding it into a known format that can later be retrieved. There are many ways to encode data, but largely these are broken into human-readable and binary formats.
CSVs, JSON, XML, and YAML are all human-readable serialization formats. Conversely, many binary formats exist, such as Parquet, Avro, and Protocol Buffers. Binary formats trade reduced readability for improved performance and size.
In the end, any data that needs to be persisted outside of a computer’s memory requires some type of serialization.
Let’s look at how serialization varies across both Rust and Python.
Python
In Python, we can serialize nearly any arbitrary data structure to JSON
using the json module.
In [1]: import json
In [2]: my_obj = [{'a': 1, 'b': None}, "foo", "bar", ("baz", "baz")]
In [3]: json.dumps(my_obj)
Out[3]: '[{"a": 1, "b": null}, "foo", "bar", ["baz", "baz"]]'
Here’s the updated project code that serializes the response from the OpenWeather API.
import os
import sys
import requests
API_KEY = os.getenv("OWM_APPID")
def get_air_pollution(lat, lon):
url = f"http://api.openweathermap.org/data/2.5/air_pollution?lat={lat}&lon={lon}&appid={API_KEY}"
body = requests.get(url).json()
return body
def parse_air_pollution(body):
aqi = body["list"][0]["main"]["aqi"]
components = body["list"][0]["components"]
return (aqi, components)
if __name__ == "__main__":
usage = f"Usage: python {__file__} <lat> <lon>"
if not API_KEY:
print("Please set OWM_APPID environment variable")
sys.exit(1)
if len(sys.argv) != 3:
print(usage)
sys.exit(1)
lat = sys.argv[1]
lon = sys.argv[2]
body = get_air_pollution(lat, lon)
aqi, components = parse_air_pollution(body)
print(f"Air Quality Index: {aqi}")
print("Components:")
for k, v in components.items():
print(f" {k}: {v}")
There are a few key things to note here.
First, we’re assuming the request was successful, that there is a JSON response body, and that it can parse correctly. If any of these assumptions are incorrect an exception will be raised, and we have no obvious way of knowing what these exceptions are or which method might raise one.
def parse_air_pollution(body):
aqi = body["list"][0]["main"]["aqi"]
components = body["list"][0]["components"]
return (aqi, components)
When parsing the response, we slice into the response body to get
various components. We’re explicitly fetching keys from a dictionary under
the assumption that the payload is properly formed. There are safer dictionary
methods to use, such as .get() which will return None if the key is missing
rather than an exception, but in our case an Exception is warranted since we
can’t do anything with the data if it’s missing.
We also haven’t explicitly typed the response from the API. This is something
we can do with mypy or other tools like pydantic, but the Python interpreter
itself has no type-guarantees.
Let’s look at how we might do this in Rust.
Rust
In Rust, we’ll need to install the serde crate as well as the json feature
for reqwest.
cargo add serde --features derive
cargo add serde_json
cargo add reqwest --features json
Because Rust is a typed language, we will define the struct that represents the data we expect. The API response looks like the following:
{
"coord": {
"lon": -122.5889,
"lat": 37.9871
},
"list": [
{
"main": {
"aqi": 2
},
"components": {
"co": 168.56,
"no": 0.14,
"no2": 0.75,
"o3": 80.11,
"so2": 0.7,
"pm2_5": 3.48,
"pm10": 5.58,
"nh3": 0
},
"dt": 1687308878
}
]
}
We can define a struct that represents this data as follows:
#![allow(unused)]
fn main() {
use serde::{Deserialize, Serialize};
#[derive(Debug, Serialize, Deserialize)]
pub struct AirPollution {
pub coord: Coord,
pub list: Vec<List>,
}
#[derive(Debug, Serialize, Deserialize)]
pub struct Coord {
pub lon: f32,
pub lat: f32,
}
#[derive(Debug, Serialize, Deserialize)]
pub struct List {
pub main: Main,
pub components: Components,
pub dt: usize,
}
#[derive(Debug, Serialize, Deserialize)]
pub struct Main {
pub aqi: u8,
}
#[derive(Debug, Serialize, Deserialize)]
pub struct Components {
pub co: f32,
pub no: f32,
pub no2: f32,
pub o3: f32,
pub so2: f32,
pub pm2_5: f32,
pub pm10: f32,
pub nh3: f32,
}
}
As you can see, the struct mirrors the underlying JSON structure. The serde
crate gives us a lot of flexibility here, in particular the section on
Attributes and the
Examples are worth spending some time on.
The reqwest crate also provides a json method that will automatically
deserialize the response body into a struct.
#![allow(unused)]
fn main() {
pub fn get_air_pollution(lat: f32, lon: f32) -> AirPollution {
let api_key = std::env::var("OWM_APPID").expect(
"Environment Variable OWM_APPID not set. Please set it to your
OpenWeatherMap API key. https://home.openweathermap.org/api_keys",
);
let url = format!(
"http://api.openweathermap.org/data/2.5/air_pollution?lat={}&lon={}&appid={}",
lat, lon, api_key
);
reqwest::blocking::get(url)
.expect("request failed")
.json()
.expect("json failed")
}
Our function now returns an AirPollution struct, instead of a String, and
reqwest’s json method will automatically deserialize the response body
to the correct type.
Rust uses type inference to reduce the amount of syntax required. While function parameters and signatures always require types, local variables can usually be inferred by the compiler.
Let’s look at how returning a typed Struct changes how we interact with the data
#![allow(unused)]
fn main() {
pub fn parse_air_pollution(body: &AirPollution) -> (&Main, &Components) {
let main = &body.list[0].main;
let components = &body.list[0].components;
(main, components)
}
}
We can access the underlying fields in the struct directly. Unlike a Python dictionary, the compiler will ensure that the fields we’re accessing exist.
If we add a missing field, for example:
#![allow(unused)]
fn main() {
let foo = &body.list[0].foo;
}
And run cargo check we’ll get the following error:
error[E0609]: no field `foo` on type `List`
--> src/bin/ch4.rs:65:29
|
65 | let foo = &body.list[0].foo;
| ^^^ unknown field
|
= note: available fields are: `main`, `components`, `dt`
For more information about this error, try `rustc --explain E0609`.
error: could not compile `wxrs` (bin "ch4") due to previous error
Compare to Python where we’d only get a run-time error if we tried to
access a missing field, unless we opt-in to type hints using mypy.
Here’s the full Rust code for reference
use serde::{Deserialize, Serialize};
#[derive(Debug, Serialize, Deserialize)]
pub struct AirPollution {
pub coord: Coord,
pub list: Vec<List>,
}
#[derive(Debug, Serialize, Deserialize)]
pub struct Coord {
pub lon: f32,
pub lat: f32,
}
#[derive(Debug, Serialize, Deserialize)]
pub struct List {
pub main: Main,
pub components: Components,
pub dt: usize,
}
#[derive(Debug, Serialize, Deserialize)]
pub struct Main {
pub aqi: u8,
}
#[derive(Debug, Serialize, Deserialize)]
pub struct Components {
pub co: f32,
pub no: f32,
pub no2: f32,
pub o3: f32,
pub so2: f32,
pub pm2_5: f32,
pub pm10: f32,
pub nh3: f32,
}
pub fn get_air_pollution(lat: f32, lon: f32) -> AirPollution {
let api_key = std::env::var("OWM_APPID").expect(
"Environment Variable OWM_APPID not set. Please set it to your
OpenWeatherMap API key. https://home.openweathermap.org/api_keys",
);
let url = format!(
"http://api.openweathermap.org/data/2.5/air_pollution?lat={}&lon={}&appid={}",
lat, lon, api_key
);
reqwest::blocking::get(url)
.expect("request failed")
.json()
.expect("json failed")
}
pub fn parse_air_pollution(body: &AirPollution) -> (&Main, &Components) {
let main = &body.list[0].main;
let components = &body.list[0].components;
(main, components)
}
pub fn main() {
let usage = format!("Usage: {} [lat] [lon]", std::env::args().next().unwrap());
let lat = std::env::args()
.nth(1)
.expect(&usage)
.parse::<f32>()
.expect(&usage);
let lon = std::env::args()
.nth(2)
.expect(&usage)
.parse::<f32>()
.expect(&usage);
let body = get_air_pollution(lat, lon);
let (main, components) = parse_air_pollution(&body);
println!("Air Quality Index: {}", main.aqi);
println!("Carbon Monoxide: {} μg/m³", components.co);
println!("Nitrogen Monoxide: {} μg/m³", components.no);
println!("Nitrogen Dioxide: {} μg/m³", components.no2);
println!("Ozone: {} μg/m³", components.o3);
println!("Sulfur Dioxide: {} μg/m³", components.so2);
println!("Particulate Matter < 2.5 μm: {} μg/m³", components.pm2_5);
println!("Particulate Matter < 10 μm: {} μg/m³", components.pm10);
println!("Ammonia: {} μg/m³", components.nh3);
}
Serialization Formats
Something worth mentioning about the Rust serde crate is that it does not
come with any built-in serialization formats. Instead, it provides a framework
for serialization. We installed serde_json but there are many other formats
available, such as serde_yaml and serde_avro.
Why Bother?
You might be wondering why we’d go through the trouble of defining a struct
and serializing the response body into a struct. In Python, we avoid the
boilerplate, we access fields directly, we can throw a little type-hinting at
our code, we get to use # type: ignore freely, and if our application crashes,
well, we’ll just fix it and run it again.
You are absolutely right! This is all true. However, any seasoned Python programmer is also aware of all the ways that poorly typed code can go wrong.
If you’ve ever created a compute-intensive application that operates on many gigabytes of data, you’ve probably run into a situation where you’ve had to re-run the application because it crashed. Type-safety helps prevent these types of issues, but types also provide another nice benefit: improved performance.
The compiler can optimize code based on the types it knows about. In Python, we can use type-hints to help the compiler, but ultimately the Python interpreter is still dynamically resolving types at runtime. In Rust, the compiler knows the types at compile-time and can optimize prior to running.
What’s that little & doing?
Ah, yes, the &. Now we are getting into the heart of Rust. Let’s look
at the code for parsing air pollution again:
#![allow(unused)]
fn main() {
pub fn parse_air_pollution(body: &AirPollution) -> (&Main, &Components) {
let main = &body.list[0].main;
let components = &body.list[0].components;
(main, components)
}
}
parse_air_pollution is a function that takes a reference to an AirPollution
struct. The & is the syntax for creating a reference. In Rust, references
are a way of passing a value to a function without transferring ownership of
the value. This is a key concept in Rust, and it’s what allows Rust to
guarantee memory safety.
In Python, values are passed around as references and tracked with counters. Every object carries a reference count, which is incremented each time a new name points at it and decremented whenever one of those names goes out of scope. When the count reaches zero, the object is freed immediately. Python also ships a separate cycle-detecting garbage collector, which runs occasionally to clean up groups of objects that reference each other and so never reach a count of zero on their own.
In Rust, there is no garbage collector. Instead, the compiler keeps track of the lifetime of every variable. When a variable goes out of scope, the compiler will automatically free the memory associated with the variable.
This means that you cannot use a variable after transferring ownership. For a deeper dive into the concept of ownership, read the Rust Book.
For example, if we tried print the value of body after assigning it, the compiler would give us an error:
#![allow(unused)]
fn main() {
fn parse_air(body: AirPollution) {
let foo = body;
println!("{:?}", body);
}
}
error[E0382]: borrow of moved value: `body`
--> src/bin/ch4.rs:71:22
|
69 | fn parse_air(body: AirPollution) {
| ---- move occurs because `body` has type `AirPollution`, which does not implement the `Copy` trait
70 | let foo = body;
| ---- value moved here
71 | println!("{:?}", body);
| ^^^^ value borrowed here after move
It’s beyond the scope of this post to explain all the details of ownership and references, but it’s important to understand that Rust’s compiler is keeping track of the lifetime of every variable, and will not allow you to use a variable after it’s been moved.
Instead, we can use a reference to a variable. This keeps the underlying data in the same place in memory, but allows us to pass it to a function as a reference to the original value.
#![allow(unused)]
fn main() {
fn parse_air(body: &AirPollution) {
let foo = body;
println!("{:?}", body);
}
}
This has some really nice benefits when it comes to processing large amounts of data, as data engineers tend to do.
In Python, it’s not always clear when data is being copied, moved, or referenced. In Rust, copying code is very explicit. If we didn’t want to borrow a reference in the code above, we could also copy.
#![allow(unused)]
fn main() {
fn parse_air(body: AirPollution) {
let foo = body.clone();
println!("{:?}", body);
}
}
For the above code to work, we would also need to implement the Clone trait
for the AirPollution struct and all of its fields:
#![allow(unused)]
fn main() {
#[derive(Debug, Clone, Deserialize)]
pub struct AirPollution {
...
}
Understanding ownership, references, and borrowing can be an uphill battle for new Rust programmers who are used to dynamically-typed languages, but with time and patience, it will come to you too.
Performance
To benchmark our code, we’re going to change our code to fetch an entire forecast rather than a single day, increasing the payload from 0.5kb to about 13kb.
In Python, we change the url and then iterate over every element in the list provided.
def get_air_pollution(lat, lon):
url = f"http://api.openweathermap.org/data/2.5/air_pollution/forecast?lat={lat}&lon={lon}&appid={API_KEY}"
body = requests.get(url).json()
return body
def parse_air_pollution(body):
res = []
print(body)
for row in body["list"]:
res.append((row["main"]["aqi"], row["components"], row["dt"]))
return res
def print_air_pollution(main, components, dt):
print("---")
print(f"Air pollution forecast for {dt}")
print(f"Air quality index: {main}")
print("Components:")
for k, v in components.items():
print(f" {k}: {v}")
In Rust, we also change the url and use the common iter().map().collect()
pattern.
#![allow(unused)]
fn main() {
pub fn get_air_pollution(lat: f32, lon: f32) -> AirPollution {
let api_key = std::env::var("OWM_APPID").expect(
"Environment Variable OWM_APPID not set. Please set it to your
OpenWeatherMap API key. https://home.openweathermap.org/api_keys",
);
let url = format!(
"http://api.openweathermap.org/data/2.5/air_pollution/forecast?lat={}&lon={}&appid={}",
lat, lon, api_key
);
reqwest::blocking::get(url)
.expect("request failed")
.json()
.expect("json failed")
}
pub fn parse_air_pollution(body: AirPollution) -> Vec<(Main, Components, usize)> {
body.list
.iter()
.map(|x| (x.main, x.components, x.dt))
.collect()
}
}
Here are the results of the benchmarks:
| Command | Mean [ms] | Min [ms] | Max [ms] | Relative |
|---|---|---|---|---|
../wxrs/target/release/ch4_benchmark 30 -140 | 112.9 ± 10.5 | 96.7 | 141.0 | 1.00 |
../wxpy/.venv/bin/python ../wxpy/wxpy/ch4/serialized_benchmark.py 30 -140 | 149.7 ± 8.6 | 139.0 | 172.4 | 1.33 ± 0.14 |
Again we see Rust come out ahead, this time by about 1.3x. As in the last chapter the margin is narrower than you might expect, and for the same reason: both programs spend most of their time waiting on the network, and most of what is left is Python’s startup. Deserializing 13kb of JSON barely registers against either.
To actually measure the parsing we need to get the network out of the way.
Offline Benchmarks
Benchmarking against a network connection can be a bit iffy. It also makes it hard to test larger and larger payloads, so we’ll create a large payload file and use that for an offline benchmark.
I’ve created a 9mb JSON file that mirrors the payload from the OpenWeather
API, and created offline versions of the Rust and Python code to read from
a local file. The code for both can be found in the sample repository under
wxpy/wxpy/ch4/serialized_offline_benchmark.py and wxrs/src/bin/ch4_offline_benchmark.rs.
Here are the results of the offline benchmarks:
| Command | Mean [ms] | Min [ms] | Max [ms] | Relative |
|---|---|---|---|---|
../wxrs/target/release/ch4_offline_benchmark | 17.6 ± 0.3 | 17.1 | 18.4 | 1.00 |
../wxpy/.venv/bin/python ../wxpy/wxpy/ch4/serialized_offline_benchmark.py | 103.5 ± 1.9 | 100.9 | 107.5 | 5.87 ± 0.14 |
With the network out of the picture and a much larger payload, the gap widens considerably: Rust is now nearly six times as fast as Python.
An aside: Rust is not automatically faster
The first time I ran this benchmark, Rust lost, and it is worth explaining why, because it is a mistake that is very easy to make.
Both programs print a line per record, and there are a lot of records. Rust’s
standard output is line buffered, which means every println! costs a write
syscall. Python’s standard output, when it is not attached to a terminal, is
block buffered, so it batches those same lines into far fewer, larger writes.
The result was a Rust program that spent most of its life in the kernel:
| Wall time | User | System | |
|---|---|---|---|
Rust, println! | 114.2 ms | 42.6 ms | 70.9 ms |
Rust, BufWriter | 16.8 ms | 15.0 ms | 1.4 ms |
Look at the system time. That is the whole story: the parsing was never the
problem. Wrapping stdout in a BufWriter is the idiomatic fix, and it is what
the code in the repository now does.
#![allow(unused)]
fn main() {
let stdout = std::io::stdout();
let mut out = BufWriter::new(stdout.lock());
}
The general lesson is one worth internalising before you rewrite anything in Rust for performance: a language that is capable of being faster will still happily let you write something slower, and the bottleneck is very often I/O rather than the computation you were focused on. Measure, and look at where the time actually goes.
Transforming Data using Polars
In this chapter, we’ll look at how to transform data using Polars in both Python and Rust.
Polars is a “blazing fast DataFrame library” available in both Python and Rust. When I first wrote this chapter it was reasonable to describe it as a faster pandas with fewer features; that framing has not aged well. Polars has since reached 1.0 on the Python side and covers most of what you would reach for pandas to do.
The Polars documentation is a great resource for getting started, and the API docs have even more detail on syntax.
One thing worth knowing up front: the two languages are on different version
numbers for the same project. The Python package is at 1.43.0 and the Rust
crate is at 0.54.4. They are not as far apart as that makes them look.
Getting the data
This chapter uses the Project FeederWatch
dataset, which is checked into the repository as a 7z archive because the
extracted CSV is about 1.4GB. Unpack it first:
# from the repo root
make data
A note on lazy vs eager
Both languages give you two ways to work: eager, where each operation runs
immediately, and lazy, where you describe the whole query and let Polars
optimise it before running anything. Lazy is where the interesting work
happens — it can push our column selection and our valid == 1 filter down
into the CSV reader, so it never materialises the columns and rows we are
going to throw away.
Both versions below use the lazy API, which keeps the comparison honest. It is also how you would write this in practice.
Let’s look at some key differences between the syntax in Python and Rust.
Python
import os
import polars as pl
script_path = os.path.dirname(os.path.realpath(__file__))
bird_path = os.path.join(script_path, "../../../lib/PFW_2016_2020_public.csv")
codes_path = os.path.join(script_path, "../../../lib/species_code.csv")
# The columns we care about, in the casing the CSV actually uses.
COLS = [
"LATITUDE",
"LONGITUDE",
"SUBNATIONAL1_CODE",
"Month",
"Day",
"Year",
"SPECIES_CODE",
"HOW_MANY",
"VALID",
]
birds = pl.scan_csv(bird_path).select([pl.col(c).alias(c.lower()) for c in COLS])
codes = pl.scan_csv(codes_path, infer_schema_length=None).select(
[
pl.col("SPECIES_CODE").alias("species_code"),
pl.col("PRIMARY_COM_NAME").alias("species_name"),
]
)
birds_df = (
birds.filter(pl.col("valid") == 1)
.group_by(["subnational1_code", "species_code"])
.agg(
[
pl.col("how_many").sum().alias("total_species"),
pl.col("how_many").count().alias("total_sightings"),
]
)
.join(codes, on="species_code", how="inner")
.sort("total_species", descending=True)
.collect()
)
print(birds_df)
The Python code is very concise. pl.scan_csv gives us a lazy frame, columns
can be selected as a list of expressions, sort takes a simple descending
argument, and nothing actually runs until the final collect().
I’ve also included an attempt at the same logic in pandas. While largely similar, there are a few differences, for example, in how we filter for valid results. Pandas has no lazy mode, so it does all of the work eagerly.
import os
import pandas as pd
script_path = os.path.dirname(os.path.realpath(__file__))
bird_path = os.path.join(script_path, "../../../lib/PFW_2016_2020_public.csv")
codes_path = os.path.join(script_path, "../../../lib/species_code.csv")
# adding usecols reducing memory usage and runtime from 13s to 7s
birds = pd.read_csv(
bird_path,
usecols=[
"LATITUDE",
"LONGITUDE",
"SUBNATIONAL1_CODE",
"Month",
"Day",
"Year",
"SPECIES_CODE",
"HOW_MANY",
"VALID",
],
).rename(columns=lambda x: x.lower())
codes = pd.read_csv(codes_path)[["SPECIES_CODE", "PRIMARY_COM_NAME"]].rename(
columns={"SPECIES_CODE": "species_code", "PRIMARY_COM_NAME": "species_name"}
)
birds = birds[
[
"latitude",
"longitude",
"subnational1_code",
"month",
"day",
"year",
"species_code",
"how_many",
"valid",
]
]
birds = birds[birds["valid"] == 1]
birds = (
birds.groupby(["subnational1_code", "species_code"])
.agg(total_species=("how_many", "sum"), total_sightings=("how_many", "count"))
.reset_index()
)
birds = pd.merge(birds, codes, on="species_code", how="inner").sort_values(
"total_species", ascending=False
)
print(birds)
Now let’s compare the above to Rust code.
Rust
use polars::prelude::*;
// Resolved at compile time relative to this crate, so the program works no
// matter which directory you run it from.
const BIRD_PATH: &str = concat!(
env!("CARGO_MANIFEST_DIR"),
"/../lib/PFW_2016_2020_public.csv"
);
const CODES_PATH: &str = concat!(env!("CARGO_MANIFEST_DIR"), "/../lib/species_code.csv");
// The columns we care about, in the casing the CSV actually uses.
const COLS: [&str; 9] = [
"LATITUDE",
"LONGITUDE",
"SUBNATIONAL1_CODE",
"Month",
"Day",
"Year",
"SPECIES_CODE",
"HOW_MANY",
"VALID",
];
fn main() -> PolarsResult<()> {
let birds = LazyCsvReader::new(BIRD_PATH.into())
.with_has_header(true)
.finish()?
.select(
COLS.iter()
.map(|name| col(*name).alias(name.to_lowercase()))
.collect::<Vec<_>>(),
);
let codes = LazyCsvReader::new(CODES_PATH.into())
.with_has_header(true)
.with_infer_schema_length(None)
.finish()?
.select([
col("SPECIES_CODE").alias("species_code"),
col("PRIMARY_COM_NAME").alias("species_name"),
]);
let joined = birds
.filter(col("valid").eq(lit(1)))
.group_by([col("subnational1_code"), col("species_code")])
.agg([
col("how_many").sum().alias("total_species"),
col("how_many").count().alias("total_sightings"),
])
.join(
codes,
[col("species_code")],
[col("species_code")],
JoinArgs::new(JoinType::Inner),
)
.sort(
["total_species"],
SortMultipleOptions::default().with_order_descending(true),
)
.collect_with_engine(Engine::Streaming)?
.unwrap_single();
println!("{}", joined);
Ok(())
}
The shape of the query is identical — scan, select, filter, group, aggregate, join, sort — but the Rust version is roughly 60% longer.
Almost all of that extra length is types and error handling rather than logic. A few things worth pointing out:
mainreturnsPolarsResult<()>, which lets us use?after every fallible call. An earlier version of this chapter was littered withunwrap; this reads better and behaves better.sorttakes aSortMultipleOptionsbuilder rather than a bare keyword argument, because Rust has no keyword arguments.LazyCsvReader::newwants aPlRefPath, not aPathBuf, so the paths go through.into().- The paths themselves are built with
concat!(env!("CARGO_MANIFEST_DIR"), ..), which resolves them at compile time relative to the crate. Python gets the same effect at runtime from__file__.
Overall the APIs are close enough that translating between them is mostly mechanical.
Benchmarks
Let’s look at some benchmarks for polars in both Python and Rust, as well as similar code in Pandas.
| Command | Mean [s] | Min [s] | Max [s] | Relative |
|---|---|---|---|---|
../wxrs/target/release/ch5 | 1.365 ± 0.023 | 1.335 | 1.400 | 3.15 ± 0.10 |
../wxpy/.venv/bin/python ../wxpy/wxpy/ch5/ch5.py | 0.433 ± 0.011 | 0.416 | 0.450 | 1.00 |
../wxpy/.venv/bin/python ../wxpy/wxpy/ch5/ch5_pandas.py | 4.107 ± 0.047 | 4.020 | 4.196 | 9.48 ± 0.26 |
The first thing to take from this table is the thing that has not changed: both Polars versions comfortably beat pandas. Rust-Polars is about 3x faster than pandas here, and Python-Polars about 9.5x. If you came to this chapter wondering whether Polars is worth adopting, that question is settled regardless of which language you write it in.
The second thing is that Python-Polars is the fastest of the three, finishing in about 0.43s against Rust’s 1.37s. That is not the result the earlier version of this chapter reported, and it is not the result I expected.
It would be easy to quietly drop that second finding. It is more interesting to sit with it, because it points at something that is true in general and easy to forget.
Both Polars versions run the same engine. Polars is written in Rust, and the Python package is a thin binding over that same Rust core. This benchmark was never really Rust versus Python. It is one build of a Rust library against another build of the same Rust library, with a small amount of Python doing the orchestration around it. The Python interpreter barely participates: look at the user time in the table above and you’ll see both Polars runs burning several CPU-seconds in parallel inside the engine.
So why is our build slower? I checked the two most obvious explanations and neither held up:
- The allocator. Polars’ own docs recommend a custom allocator and say it
can be worth up to 25%. Swapping in
mimallocchanged the runtime by less than the run-to-run noise here, so I took it back out. - The engine. Recent Polars has both an in-memory and a streaming engine,
and Python’s
collect()chooses differently than Rust’s does. Forcing the streaming engine in Rust withcollect_with_engine(Engine::Streaming)took 1.43s down to 1.34s — real, but nowhere near a 3x gap.
What is left is the build itself. The Python wheels are compiled with tuning
that a plain cargo build --release does not apply — the
Polars performance notes
recommend a nightly compiler with the simd and performant features and
RUSTFLAGS='-C target-cpu=native'. I have not chased that here, partly because
target-cpu=native produces a binary tuned to whatever machine built it, which
is at odds with pinning everything else in this repository so the numbers
reproduce.
The lesson I would take from this is the same one from the BufWriter aside in
the last chapter, one level up. Reaching for Rust does not hand you
performance. When you call into a library that is already written in Rust,
choosing Rust as your language may buy you very little — you were always
running Rust, and what actually mattered was how somebody else compiled it.
Which is worth holding next to the pandas column. The 9.5x that separates Python-Polars from pandas came from choosing a better tool. The 3x that separates it from our Rust build came from choosing a better build of the same tool. Neither of those is a fact about Python or Rust the languages, and picking the right library will usually take you further than picking the right language.
Concurrent Programming
One of Rust’s major goals as a language is to enable fearless concurrency. So much so that an entire chapter of the Rust Book is devoted to it.
In Python, concurrency is possible however we are impacted by the GIL.
What’s really fascinating (to me, anyways) is how decisions about how memory is managed in both languages has a direct impact on how concurrency is handled.
Before we dig into concurrency, let’s take a step back and talk about memory.
Memory
Every programming language stores objects in memory. Whether it’s variables, functions, or other data, we store these in memory to allow fast access to them when we need them.
How languages manage memory defines the flavor and performance characteristics of the language.
The GIL and Python’s Memory Management
In Python, the infamous Global Interpreter Lock (or GIL) exists because objects in Python are reference counted. This means that every object has a counter associated with it that is incremented as it is referenced and decremented as it is removed from scope. When an object has 0 references, it is cleared from memory, freeing up space.
Those counters are not themselves thread-safe: if two threads incremented or decremented the same count at once, the object could be freed while still in use, or leak forever. Rather than lock every object individually, CPython takes a single global lock, the GIL, which guarantees that only one thread executes Python bytecode at a time. This has the effect of serializing execution and effectively making CPU-bound Python code single-threaded, no matter how many threads you spawn.
To work around these limitations, CPU-bound Python code has to reach for
separate processes, typically via the multiprocessing module, which sidesteps
the GIL by giving each process its own interpreter. That comes with its own set
of limitations and overhead costs, since data has to be pickled and copied
between processes rather than simply shared.
About the Author
This Rust for Data book was created by me, Pedram Navid.
You can find me on Twitter @pdrmnvd
and on LinkedIn @pedramnavid
and on GitHub @pedramnavid
and on Substack @databased.
and on my website pedramnavid.com.