DEV Community

William
William

Posted on

Protobuf without `protoc`: the ultimate developer experience for Python

protoruf logo

Install protoc, generate _pb2.py files, juggle Parse and SerializeToString… Sounds familiar? With protoruf, all of that disappears. On-the-fly compilation, zero generated files, and native Pydantic integration. Here's how I swapped the complexity of google.protobuf for absolute simplicity.


If you've ever used Protobuf in Python, you know the routine:

  • Manually install protoc (and pray the version matches everywhere, especially in CI).
  • Generate Python code with protoc --python_out=., then version or regenerate that file every time the schema changes.
  • Write procedural code full of Parse, SerializeToString, MessageToJson calls, with no native data validation.

It works, but it's clunky. And every schema change means repeating the whole ritual.

Protoruf was built to fix this at the root. It's a Python library written in Rust that compiles .proto files in memory, converts data in a single line, and integrates naturally with Pydantic for validation.

The result? A radically smoother developer experience, where you can modify your Protobuf schema and keep coding immediately — without ever leaving your editor.


Before: the classic google.protobuf workflow

# Step 1 – Install protoc (a CI nightmare)
sudo apt-get install protobuf-compiler   # Linux
brew install protobuf                    # macOS
# Windows: download an .exe, add to PATH…

# Step 2 – Generate static Python code
protoc --python_out=. message.proto
# → A message_pb2.py file appears.
# You must version it or regenerate it on every build.
Enter fullscreen mode Exit fullscreen mode
# Step 3 – Write the code
from google.protobuf import json_format
from message_pb2 import Message  # The generated file

msg = Message()
json_str = '{"id": "123", "content": "Hello"}'
json_format.Parse(json_str, msg)           # Parse JSON → Protobuf

binary = msg.SerializeToString()           # Serialize

msg2 = Message()
msg2.ParseFromString(binary)               # Deserialize
result = json_format.MessageToJson(msg2)   # Back to JSON
Enter fullscreen mode Exit fullscreen mode

Pain points:

  • Dependency on protoc (manual installation, version mismatches).
  • Mandatory code generation → files to version or integrate into builds.
  • Schema change = regenerate everywhere.
  • No upfront data validation (everything is plain dict).

After: the protoruf simplicity

# A single command, end of story.
pip install protoruf
# No protoc, no external compiler.
Enter fullscreen mode Exit fullscreen mode
from protoruf import compile_proto, json_to_protobuf, protobuf_to_json

# 1. Built-in compilation, in memory (Rust)
descriptor = compile_proto("message.proto")

# 2. JSON → Protobuf in 1 line
json_str = '{"id": "123", "content": "Hello"}'
binary = json_to_protobuf(json_str, descriptor, message_type="message.Message")

# 3. Protobuf → JSON in 1 line
result = protobuf_to_json(binary, descriptor, message_type="message.Message", pretty=True)
print(result)
Enter fullscreen mode Exit fullscreen mode

What changes:

  • No more protoc — everything is bundled.
  • Zero generated files → your repo stays clean.
  • Schema update? Just re-run your script.
  • Fewer lines, less verbosity, more clarity.

Bonus: native Pydantic integration

protoruf provides direct functions to convert a Pydantic model to a Protobuf message and back. No more manual dict fiddling.

from pydantic import BaseModel
from protoruf import pydantic_to_protobuf, protobuf_to_pydantic

class Message(BaseModel):
    id: str
    content: str

msg = Message(id="123", content="Hello")

# Direct conversion, no manual JSON dump
binary = pydantic_to_protobuf(msg, descriptor, "message.Message")

# Back to a Pydantic instance
result = protobuf_to_pydantic(binary, descriptor, Message, "message.Message")
print(result.content)  # Hello
Enter fullscreen mode Exit fullscreen mode

You get automatic validation from Pydantic both on input and output, with zero extra code.


Bonus 2: Blazing fast with DescriptorCache

Because protoruf is built on Rust (prost-reflect), it is incredibly fast. But for high-throughput workloads (like processing thousands of messages in a loop), protoruf offers a DescriptorCache.

By decoding the descriptor pool once and reusing it, protoruf achieves 5x to 10x faster serialization/deserialization compared to google.protobuf.

from protoruf import compile_proto, DescriptorCache

# Decode the pool ONCE
cache = DescriptorCache(compile_proto("message.proto"))

# Reuse it everywhere (Thread-safe!)
for json_str in massive_data_stream:
    binary = cache.json_to_protobuf(json_str, "message.Message")
    # Process binary...
Enter fullscreen mode Exit fullscreen mode

Benchmarks

Transparency matters, so here are the numbers from my local machine (100,000 messages, same payload for both libraries).

Simple usage (free functions, no cache)

This is the "just call json_to_protobuf" approach — no setup, no cache. The descriptor is decoded on every call.

Operation google.protobuf protoruf (free functions)
Serialization (JSON→Proto) 1.83 s (54.7k msg/s) 2.07 s (48.2k msg/s)
Parsing (Proto→JSON) 1.62 s (61.8k msg/s) 1.88 s (53.2k msg/s)

In the simplest usage, protoruf is about 13–16% slower because it decodes the descriptor set from scratch on every call. That's the trade-off for not needing code generation upfront.

Hot loop with DescriptorCache

Now the same conversion, but with the descriptor pool decoded once outside the loop (just like you'd do in a real service).

Operation google.protobuf protoruf (DescriptorCache)
Serialization (JSON→Proto) 1.89 s (53.0k msg/s) 0.36 s (280k msg/s)
Parsing (Proto→JSON) 1.65 s (60.5k msg/s) 0.16 s (612k msg/s)

With the cache, protoruf is 5.3× faster on writes and 10.1× faster on reads. That's the power of decoding once and reusing — and it's the recommended setup for any service processing more than a handful of messages.

The best part? You get both modes with the same library. Start simple, and when you need the extra throughput, just wrap your descriptor in DescriptorCache.


Why the developer experience is transformed

Aspect google.protobuf protoruf
Installation protoc + pip pip install protoruf
Code generation Yes (_pb2.py files) None (in-memory compilation)
Schema changes Regenerate everywhere, redeploy Do nothing, just re-run
CI / CD Nightmare (protoc versioning) Zero configuration
Data validation Manual Native Pydantic integration
Lines of code ~15 lines ~5 lines
Performance (Hot loops) Baseline Up to 10x faster (with Cache)

With protoruf, the time you used to spend configuring protoc and maintaining generated files goes directly into your real code. Changing a Protobuf schema is no longer a chore — it's just another line of code.


Try it in 30 seconds

pip install protoruf
Enter fullscreen mode Exit fullscreen mode

Create a message.proto:

syntax = "proto3";
package message;
message Message {
  string id = 1;
  string content = 2;
  int32 priority = 3;
}
Enter fullscreen mode Exit fullscreen mode

Then run:

from protoruf import compile_proto, json_to_protobuf, protobuf_to_json

descriptor = compile_proto("message.proto")

pb = json_to_protobuf('{"id":"42","content":"Hello"}', descriptor, "message.Message")
print(protobuf_to_json(pb, descriptor, "message.Message", pretty=True))
Enter fullscreen mode Exit fullscreen mode

No hidden steps, no stray files. You code, it compiles.


In summary

Protoruf is the choice of productivity and peace of mind for all your Protobuf exchanges in Python.

No more protoc, no more generated files, no more verbose gymnastics.

You keep the power of Protobuf, with a Pythonic API and instant Pydantic integration.



🔗 Links:

Feedback, issues, and PRs are more than welcome! Let me know in the comments what you think.

Top comments (1)

Collapse
 
jeanne_bjar profile image
Jeanne Béjar •

Very useful concept! It actually speeds up a lot my backend communications, thank you for the lib!