Install protoc, generate _pb2.py files, juggle Parse and SerializeToString… Sounds familiar? With protoruf, all of that disappears. On-the-fly compilation, zero generated files, and native Pydantic integration. Here's how I swapped the complexity of google.protobuf for absolute simplicity.
If you've ever used Protobuf in Python, you know the routine:
- Manually install
protoc(and pray the version matches everywhere, especially in CI). - Generate Python code with
protoc --python_out=., then version or regenerate that file every time the schema changes. - Write procedural code full of
Parse,SerializeToString,MessageToJsoncalls, with no native data validation.
It works, but it's clunky. And every schema change means repeating the whole ritual.
Protoruf was built to fix this at the root. It's a Python library written in Rust that compiles .proto files in memory, converts data in a single line, and integrates naturally with Pydantic for validation.
The result? A radically smoother developer experience, where you can modify your Protobuf schema and keep coding immediately — without ever leaving your editor.
Before: the classic google.protobuf workflow
# Step 1 – Install protoc (a CI nightmare)
sudo apt-get install protobuf-compiler # Linux
brew install protobuf # macOS
# Windows: download an .exe, add to PATH…
# Step 2 – Generate static Python code
protoc --python_out=. message.proto
# → A message_pb2.py file appears.
# You must version it or regenerate it on every build.
# Step 3 – Write the code
from google.protobuf import json_format
from message_pb2 import Message # The generated file
msg = Message()
json_str = '{"id": "123", "content": "Hello"}'
json_format.Parse(json_str, msg) # Parse JSON → Protobuf
binary = msg.SerializeToString() # Serialize
msg2 = Message()
msg2.ParseFromString(binary) # Deserialize
result = json_format.MessageToJson(msg2) # Back to JSON
Pain points:
- Dependency on
protoc(manual installation, version mismatches). - Mandatory code generation → files to version or integrate into builds.
- Schema change = regenerate everywhere.
- No upfront data validation (everything is plain
dict).
After: the protoruf simplicity
# A single command, end of story.
pip install protoruf
# No protoc, no external compiler.
from protoruf import compile_proto, json_to_protobuf, protobuf_to_json
# 1. Built-in compilation, in memory (Rust)
descriptor = compile_proto("message.proto")
# 2. JSON → Protobuf in 1 line
json_str = '{"id": "123", "content": "Hello"}'
binary = json_to_protobuf(json_str, descriptor, message_type="message.Message")
# 3. Protobuf → JSON in 1 line
result = protobuf_to_json(binary, descriptor, message_type="message.Message", pretty=True)
print(result)
What changes:
- No more
protoc— everything is bundled. - Zero generated files → your repo stays clean.
- Schema update? Just re-run your script.
- Fewer lines, less verbosity, more clarity.
Bonus: native Pydantic integration
protoruf provides direct functions to convert a Pydantic model to a Protobuf message and back. No more manual dict fiddling.
from pydantic import BaseModel
from protoruf import pydantic_to_protobuf, protobuf_to_pydantic
class Message(BaseModel):
id: str
content: str
msg = Message(id="123", content="Hello")
# Direct conversion, no manual JSON dump
binary = pydantic_to_protobuf(msg, descriptor, "message.Message")
# Back to a Pydantic instance
result = protobuf_to_pydantic(binary, descriptor, Message, "message.Message")
print(result.content) # Hello
You get automatic validation from Pydantic both on input and output, with zero extra code.
Bonus 2: Blazing fast with DescriptorCache
Because protoruf is built on Rust (prost-reflect), it is incredibly fast. But for high-throughput workloads (like processing thousands of messages in a loop), protoruf offers a DescriptorCache.
By decoding the descriptor pool once and reusing it, protoruf achieves 5x to 10x faster serialization/deserialization compared to google.protobuf.
from protoruf import compile_proto, DescriptorCache
# Decode the pool ONCE
cache = DescriptorCache(compile_proto("message.proto"))
# Reuse it everywhere (Thread-safe!)
for json_str in massive_data_stream:
binary = cache.json_to_protobuf(json_str, "message.Message")
# Process binary...
Benchmarks
Transparency matters, so here are the numbers from my local machine (100,000 messages, same payload for both libraries).
Simple usage (free functions, no cache)
This is the "just call json_to_protobuf" approach — no setup, no cache. The descriptor is decoded on every call.
| Operation | google.protobuf |
protoruf (free functions) |
|---|---|---|
| Serialization (JSON→Proto) | 1.83 s (54.7k msg/s) | 2.07 s (48.2k msg/s) |
| Parsing (Proto→JSON) | 1.62 s (61.8k msg/s) | 1.88 s (53.2k msg/s) |
In the simplest usage, protoruf is about 13–16% slower because it decodes the descriptor set from scratch on every call. That's the trade-off for not needing code generation upfront.
Hot loop with DescriptorCache
Now the same conversion, but with the descriptor pool decoded once outside the loop (just like you'd do in a real service).
| Operation | google.protobuf |
protoruf (DescriptorCache) |
|---|---|---|
| Serialization (JSON→Proto) | 1.89 s (53.0k msg/s) | 0.36 s (280k msg/s) |
| Parsing (Proto→JSON) | 1.65 s (60.5k msg/s) | 0.16 s (612k msg/s) |
With the cache, protoruf is 5.3× faster on writes and 10.1× faster on reads. That's the power of decoding once and reusing — and it's the recommended setup for any service processing more than a handful of messages.
The best part? You get both modes with the same library. Start simple, and when you need the extra throughput, just wrap your descriptor in DescriptorCache.
Why the developer experience is transformed
| Aspect | google.protobuf |
protoruf |
|---|---|---|
| Installation |
protoc + pip
|
pip install protoruf |
| Code generation | Yes (_pb2.py files) |
None (in-memory compilation) |
| Schema changes | Regenerate everywhere, redeploy | Do nothing, just re-run |
| CI / CD | Nightmare (protoc versioning) |
Zero configuration |
| Data validation | Manual | Native Pydantic integration |
| Lines of code | ~15 lines | ~5 lines |
| Performance (Hot loops) | Baseline | Up to 10x faster (with Cache) |
With protoruf, the time you used to spend configuring
protocand maintaining generated files goes directly into your real code. Changing a Protobuf schema is no longer a chore — it's just another line of code.
Try it in 30 seconds
pip install protoruf
Create a message.proto:
syntax = "proto3";
package message;
message Message {
string id = 1;
string content = 2;
int32 priority = 3;
}
Then run:
from protoruf import compile_proto, json_to_protobuf, protobuf_to_json
descriptor = compile_proto("message.proto")
pb = json_to_protobuf('{"id":"42","content":"Hello"}', descriptor, "message.Message")
print(protobuf_to_json(pb, descriptor, "message.Message", pretty=True))
No hidden steps, no stray files. You code, it compiles.
In summary
Protoruf is the choice of productivity and peace of mind for all your Protobuf exchanges in Python.
No more protoc, no more generated files, no more verbose gymnastics.
You keep the power of Protobuf, with a Pythonic API and instant Pydantic integration.

Top comments (1)
Very useful concept! It actually speeds up a lot my backend communications, thank you for the lib!