DEV Community

Ciphemic academia
Ciphemic academia

Posted on Originally published at ciphemicacademia.in

JSON vs. Protocol Buffers vs. Avro: 60% Less Latency, 80% Smaller Payloads, One Serialization Decision

JSON vs. Protocol Buffers vs. Avro: 60% Less Latency, 80% Smaller Payloads, One Serialization Decision

LinkedIn's engineering team moved their internal microservice traffic off JSON and onto Protocol Buffers, and measured a 60% drop in service latency, without touching their underlying REST framework at all, just swapping the format carrying the data. Atlassian ran the same kind of migration on part of Jira's backend and cut that cluster's data size by 80% and its CPU usage by 75%, enough to shrink the cluster itself by more than half. Neither team changed what their services did. They changed how the bytes on the wire were shaped, and that single decision showed up directly on their infrastructure bill.

This is what a serialization format actually is: the rules for turning your in-memory data into bytes that can cross a network or sit on disk, and back again. JSON, Protocol Buffers (protobuf), and Avro all do this, but they make genuinely different trade-offs, and picking the wrong one for your situation is exactly how a team ends up re-doing this kind of migration two years in. This guide breaks down what each format actually does under the hood, where the real numbers come from, and how to decide before you've shipped a year of JSON APIs you'll later have to migrate off.

If you're newer to how data moves through a system at all, our Binary & Encoding course covers the fundamentals this comparison builds on.

This post originally appeared on the Ciphemic Academia blog.

The Short Version

  • JSON is a human-readable, text-based format with no required schema. Universally supported, easy to debug, and the default most APIs still start with.
  • Protocol Buffers is a binary, schema-first format from Google. Smaller and faster than JSON, with generated, type-safe code in whatever language you need.
  • Avro is a binary, schema-first format from the Apache Hadoop ecosystem, built around strong schema evolution and tight integration with big-data pipelines like Kafka.

If you want one default: start with JSON. Move to Protobuf when you're optimizing internal service-to-service traffic for size and speed. Move to Avro specifically when you're in a data pipeline or event-streaming context where schema evolution across many producers and consumers is the real problem.

What All Three Are Actually Doing

Before the differences, the shared job, since this is what every serialization format exists to do:

  • Encoding: turning an in-memory object (a struct, a dict, a record) into a sequence of bytes
  • Decoding: turning those bytes back into a usable object on the other end
  • Defining structure: some way of saying "this data has a name field that's a string and an age field that's a number," whether that's enforced strictly or left loose
  • Handling change over time: every real system eventually adds a field, removes one, or renames one, and each format has a different answer for what happens to old data and old readers when that happens

The size and speed numbers above come directly from how differently each format approaches that last point, structure, and that first point, encoding.

Same data, three encodings: JSON stays biggest, Protobuf and Avro shrink it

JSON

What it actually is: a text-based format where data is represented as nested objects, arrays, strings, numbers, and booleans, written in a syntax most developers already know from JavaScript. There's no required schema, any JSON parser can read any valid JSON, and every field name is spelled out, in full, every single time it appears.

A small example:

{
  "id": 482,
  "customer_name": "Alex Chen",
  "total": 129.99,
  "items": [{"sku": "A1", "qty": 2}]
}
Enter fullscreen mode Exit fullscreen mode

Notice "customer_name" and "sku" and "qty" are written out as text, every time, in every single record. That repetition is exactly what LinkedIn's and Atlassian's teams were measuring when they counted bytes.

Where it shines:

  • Universal support: every language, every tool, every browser can read and write JSON with zero setup
  • Human-readable: you can open a JSON payload in a text editor and understand it immediately, which makes debugging dramatically faster
  • No schema required: you can start sending data immediately without agreeing on a formal structure in advance
  • Massive tooling ecosystem: logging, API testing tools, documentation generators, and nearly every web framework assume JSON by default

Where it struggles:

  • Verbose by design: repeating every field name in every record is exactly where JSON's size disadvantage comes from, and it compounds at scale
  • No enforced structure: without a schema, nothing stops a producer from silently sending "total": "129.99" as a string instead of a number, and that mismatch surfaces as a runtime bug, not a build-time error
  • Slower to parse: text parsing is inherently more work for a CPU than reading dense binary fields, which is where real latency differences like LinkedIn's come from
  • No native schema evolution strategy: adding or changing fields is informal, and nothing enforces backward or forward compatibility between old and new versions of your data

Who this suits: public APIs, anything where human-readability and zero-setup compatibility matter more than raw size or speed, and the correct default for most new projects.

Protocol Buffers

What it actually is: a binary, schema-first format where you define your data's structure in a .proto file, then generate code in your target language from it. Fields are identified by small numeric tags instead of full names, so the wire format only sends 1: 482 instead of "id": 482, which is most of where the size savings come from.

A small example (a .proto schema):

message Order {
  int32 id = 1;
  string customer_name = 2;
  double total = 3;
  repeated Item items = 4;
}
Enter fullscreen mode Exit fullscreen mode

The numbers (= 1, = 2, = 3) are the actual field tags sent on the wire, not the field names. That's the core trick: the schema lives in the .proto file, shared and version-controlled, not repeated in every message.

Where it shines:

  • Smaller payloads: encoding Avro and Protobuf against compressed JSON in one published benchmark across varied real-world data showed Protobuf cutting an additional 67% to 80% off the already-compressed JSON size, which is the same category of saving behind Atlassian's 80% figure
  • Faster serialization: binary encoding with numeric tags is cheaper for a CPU to produce and parse than text parsing, directly contributing to the latency gains LinkedIn measured
  • Strong, enforced contracts: the .proto file generates type-safe client and server code, catching structural mismatches at compile time instead of in production
  • Built-in schema evolution rules: adding a new field with a new tag number is safe by design, old readers simply ignore fields they don't recognize

Where it struggles:

  • Not human-readable: you can't open a Protobuf payload in a text editor and make sense of it; you need the schema and the right tooling
  • Requires a build step: every service needs the generated code regenerated whenever the .proto file changes, which is real workflow overhead compared to JSON's zero-setup nature
  • Smaller talent pool: fewer developers have hands-on Protobuf experience than JSON, which is a real onboarding cost for a team
  • Awkward for ad-hoc, human-facing debugging: inspecting a raw Protobuf message during an incident takes more steps than just reading JSON in a browser

Who this suits: internal service-to-service communication, performance-sensitive APIs, gRPC-based microservices, and any system where LinkedIn's or Atlassian's exact problem, too much JSON overhead at real scale, has already shown up as a real bottleneck.

Avro

What it actually is: a binary, schema-first format from the Apache Hadoop ecosystem, where the schema is defined in JSON (yes, Avro's own schema definition is JSON) and typically travels alongside the data itself, either embedded in the file or stored centrally in a schema registry. This makes Avro especially well-suited to systems like Kafka, where many independent producers and consumers need to agree on evolving data shapes over time.

A small example (an Avro schema, defined in JSON):

{
  "type": "record",
  "name": "Order",
  "fields": [
    {"name": "id", "type": "int"},
    {"name": "customer_name", "type": "string"},
    {"name": "total", "type": "double"}
  ]
}
Enter fullscreen mode Exit fullscreen mode

Unlike Protobuf, there are no manually-assigned numeric tags. Avro instead relies on the schema itself, often paired with a schema registry, to resolve differences between what a writer encoded and what a reader expects.

Where it shines:

  • Strong schema evolution model: Avro's reader/writer schema resolution is specifically designed for the case where producers and consumers are updated independently and at different times, which is extremely common in streaming pipelines
  • Compact size: since the schema isn't repeated in every record, Avro files are genuinely small; one published benchmark measured Avro at 100% size reduction against uncompressed JSON in its best cases and still meaningfully ahead of compressed JSON in realistic ones
  • Deep ecosystem fit: Avro is the default serialization format in much of the Hadoop and Kafka world, with first-class support in tools like Kafka Connect and Confluent's Schema Registry
  • Splittable for big-data processing: Avro files work naturally with distributed processing frameworks that need to split large files into chunks

Where it struggles:

  • Schema must be available to decode: unlike Protobuf's self-describing numeric tags, you genuinely cannot read Avro data without its schema, which makes a schema registry close to mandatory in production
  • Less common outside data engineering: Avro is comparatively rare in typical web backend or mobile API contexts, which narrows where this specific skill transfers
  • Real-world performance can disappoint without the right library: one published .NET benchmark found a specific Avro library performing worse than plain JSON serialization, a reminder that library quality, not just format choice, affects your real numbers
  • Smaller general tooling ecosystem outside its niche: compared to Protobuf's broad gRPC ecosystem, Avro's tooling is concentrated specifically around data pipelines

Who this suits: Kafka-based event streaming, Hadoop-ecosystem data pipelines, and any system where many independent teams produce and consume evolving data over a long time horizon.

Side-by-Side Comparison

JSON Protocol Buffers Avro
Format type Text Binary Binary
Schema None required Required, in .proto files Required, defined in JSON
Human-readable Yes No No
Typical size vs. JSON Baseline Significantly smaller Significantly smaller
Schema evolution Informal, unenforced Strong, via numeric tags Strong, via reader/writer resolution
Needs schema to decode No No (tags are self-describing) Yes, close to mandatory
Best ecosystem fit Web APIs, public APIs gRPC, internal microservices Kafka, Hadoop, data pipelines
Setup overhead None Build step, code generation Schema registry, typically

How This Fits a Career Path

  • Backend Engineer (general): JSON fundamentals are non-negotiable, since nearly every web API still uses it, and understanding its real costs is what makes the case for anything else legible
  • Microservices or platform engineer: Protobuf becomes directly relevant once internal API performance is a real, measured problem, not a hypothetical one; our API Design course covers where this decision fits alongside REST and gRPC more broadly
  • Data Engineer: Avro is close to a baseline expectation in Kafka and Hadoop-adjacent roles, and understanding schema registries and reader/writer resolution is a real, frequently-tested skill
  • System design interviews: be ready to justify a serialization choice with the same kind of reasoning LinkedIn and Atlassian used, a specific, measured problem, not a format preference

How to Choose Without Overthinking It

  1. Start with who's reading your data. Public clients, browsers, or third parties who need to debug payloads easily? JSON. Internal services you control end to end? Protobuf or Avro become worth considering.
  2. Look for an actual, measured bottleneck before migrating. LinkedIn and Atlassian didn't start with Protobuf, they measured a real cost in JSON first. Don't rewrite a working API because a format sounds more advanced.
  3. If you're in a streaming or big-data pipeline, lean Avro. Its schema evolution model is specifically built for many independent producers and consumers changing over time, which is exactly the Kafka and Hadoop situation.
  4. If you're optimizing internal RPC calls, lean Protobuf. Its tooling, gRPC integration, and broader adoption outside data engineering make it the more generally useful binary format to learn first.

A note on honesty: the specific percentages above came from real, published migrations and benchmarks, but your numbers will depend heavily on your actual data shape, field count, and library choice, the .NET Avro benchmark above found a specific library that underperformed plain JSON. Measure your own system before committing to a number you read in someone else's blog post, including this one.

Common Mistakes When Learning Serialization

  • Reaching for a binary format before you have a measured problem. Protobuf and Avro both add real setup cost; that cost needs to be paying for something concrete.
  • Treating "schema-first" as optional busywork. Skipping proper schema evolution planning is exactly how teams end up with the kind of incompatible-message incidents these formats exist to prevent.
  • Assuming all binary formats perform the same. As the .NET benchmark above shows, library quality within a format matters as much as the format itself, benchmark your actual stack, not someone else's blog post.
  • Ignoring human-debuggability costs. Moving everything to Protobuf or Avro without keeping any human-readable path (even just for internal debugging) makes incident response measurably harder.
  • Never revisiting the decision. A format chosen for a small early-stage service can become a real bottleneck at scale, as both LinkedIn's and Atlassian's stories show, revisit the decision when the data actually tells you to.

Frequently Asked Questions

Should a beginner learn JSON, Protobuf, or Avro first?

JSON. It underlies most web development you'll touch early on, requires no extra tooling, and understanding its real limitations is exactly what makes Protobuf and Avro make sense later.

Is Protobuf always faster than JSON?

In the vast majority of measured cases, yes, both in size and in CPU cost to encode and decode, which is exactly what drove LinkedIn's 60% latency improvement. But the margin depends on your specific data shape and libraries, so measuring your own case is still worth doing before committing.

Why would I choose Avro over Protobuf if they're both binary and schema-first?

Avro's reader/writer schema resolution is specifically built for situations where many independent producers and consumers evolve their data shape over time without perfect coordination, which is the normal state of a large Kafka deployment. Protobuf's numeric-tag approach to evolution is simpler but assumes slightly tighter coordination through shared .proto files.

Do I need a schema registry if I use Avro?

In production, close to always yes. Since Avro data generally cannot be decoded without its schema, a central schema registry (like Confluent's) is the standard way teams keep producers and consumers in sync.

Can I mix formats in one system?

Yes, and real systems often do, JSON facing external clients, Protobuf for internal gRPC services, Avro inside a Kafka pipeline. These aren't mutually exclusive choices for an entire company.

How much engineering effort does a JSON-to-Protobuf migration actually take?

Atlassian's and LinkedIn's write-ups both describe incremental, no-downtime migrations run over real engineering time, not a weekend project, typically involving dual-format support during the transition. Budget for genuine migration work, not a quick swap.

Measure Before You Migrate

The real lesson from LinkedIn's and Atlassian's numbers isn't "always use Protobuf," it's that they measured a specific, real cost before changing anything. Explore the Binary & Encoding course to understand what's actually happening on the wire in your own systems, so your next format decision is based on your own numbers, not someone else's blog post.

Top comments (0)