01 article

Rust Made One Function 100x Faster and the App Just 15%

FFI refactoring means rewriting a single hot function in Rust and importing it into your Python process via PyO3. One function got 100x faster, but the whole app only moved about 15 percent. The gap is where the real story lives.

Rust Made One Function 100x Faster and the App Just 15%

Rust Made One Function 100x Faster and the App Just 15%

Every performance conversation I end up in goes the same way. Someone says we should rewrite the hot path in Rust, and everyone else quietly pictures a six month migration, a team that has to learn a new language, and a folder full of C style FFI glue that nobody wants to own. The usual escape is to split the thing into microservices and push the slow part into a faster service. That works. It also buys you network latency, a new deploy surface, and the slowest possible way to fall back when the two halves disagree about what a number should be.

There is a quieter option that gets less airtime. You rewrite one function. Just the function. You compile it to a shared library, import it into your existing Python process like any other module, and leave the rest of the system alone. I have been calling this FFI refactoring, after the talk Lily Mara gave on the approach, and it sits in the gap between doing nothing and rewriting the monolith.

The idea is to rewrite one function, not the app

The whole trick leans on PyO3, the crate that lets Rust talk to the Python interpreter. You mark a function with the #[pyfunction] attribute, and PyO3 generates a Python side wrapper under the same name, so the rest of the codebase calls it exactly as before. Maturin, the build tool the PyO3 maintainers ship, compiles the crate into a C extension module that Python loads on import.

A minimal version looks like this:

#[pyfunction]
fn compute_stats(values: Vec) -> PyStats {
    // mean, std, quartiles in Rust
    PyStats { mean, std, q1, q3 }
}

That is about the entire integration surface. No gRPC client, no service mesh, no new environment variable pointing at an endpoint. The function lives in the same process and the same heap as the code that calls it, and that single fact does most of the heavy lifting later in the story.

What the compiler forces you to see

Here is the part I did not expect to enjoy. Rust's Option type makes you handle the empty case at the boundary, out in the open. In Python, a stats function that can return nothing will just return None, and every caller has to remember to check, and some of them will not. In Rust you cannot pull a value out of an Option without saying what happens when there is no value there. When that function is exposed through #[pyfunction], the compiler will not even let you return an Option directly, because Python has no idea how to unwrap it. You have to decide, on purpose, whether an empty result is an error.

That decision used to be a code review argument. Now it is a compile error you fix before the function exists.

The numbers, and why they lie a little

Run the same stats over a large input and the Rust version comes out roughly a hundred times faster than the Python one. That is the number everyone quotes. The honest version is a bit less flattering.

The Python function took about 86 microseconds per call. That is fine for a single call, but it runs at the edge of every API handler, so the CPU cost compounds across the whole fleet. Rewrite it and you cut that cost out. So far so good, a hundred times faster.

Now run a macrobenchmark of the whole application with the work tool. It lands around fifteen percent faster. Not a hundred. The function got a hundred times faster; the app is only fifteen percent. The gap is the entire story. One hot function is rarely the whole bottleneck, and the FFI call still carries overhead, small as it is. You will not put this on a conference slide. Fifteen percent off the p95 of your main endpoint is still real money, though.

Where this beats a microservices rewrite

The function level approach has a property a service boundary does not. When I rewrote the stats function, the Rust numbers and the Python numbers did not line up exactly. The quartiles came back slightly different, which is what happens when you swap one ecosystem's math for another's.

Because the Rust function and the Python function share a process, I could keep the quartile calculation in Python and move the rest to Rust. The two coexist, and falling back is a function call, not a network round trip. In a microservices rewrite, that same mismatch discovered months in would be a much uglier conversation, because the old implementation sits on the other side of a wire and keeping the old one stops being something you can do cheaply.

The tooling is finally friendly

PyO3 wants Rust 1.83 or newer, and the free threading Python builds have changed a few details about avoiding copies across the boundary, but day to day it is easy enough now. Maturin handles the build. PyCharm 2026.2 ships an official Rust plugin, so you write, compile, run, and debug the Python and the Rust in one window instead of juggling a second terminal. That sounds small. Debugging across an FFI boundary was the historical dealbreaker for a lot of teams, and it just became normal.

When it is worth it

This is not for your Go codebase. Go and Rust are close enough in performance that the FFI overhead eats the win, and it is not worth doing at all. It is for the Python bottleneck where one function is genuinely hot and everything around it is fine. Find the function that shows up red in the profiler, move that one function, keep the fallback, and measure the app rather than the function. The function will look heroic. The app is the number you actually get paid on.

Comments