Protobuf Everything? When Your Type Safety Becomes a Straightjacket

Protobuf Everything? When Your Type Safety Becomes a Straightjacket

The binary-first movement is pushing Protobuf into every layer of the stack. Here’s where it genuinely helps, and where it’s just trading one set of problems for a worse set.

There’s a particular kind of developer who discovers Protobuf and suddenly sees binary serialization everywhere. The REST endpoint that returns a list of users? Binary. The auth token? Binary. The database column that could be a perfectly readable JSON blob? You guessed it, binary.

This isn’t hypothetical. The “Protobuf Everywhere” philosophy has real advocates, and the Lupyd team recently published a detailed breakdown of exactly this approach across their web, mobile, and database layers. They’re using Protocol Buffers for client HTTP endpoints, authentication tokens, database storage, and cryptographic workflows, all without touching gRPC.

On paper, it’s a compelling story. In practice, it’s a textbook case of a tool’s strengths being applied where they help, and also where they hurt.

First, the Legit Case for Protobuf-Everywhere

Before we dunk on the binary-first crowd, credit where it’s due. There are genuinely excellent reasons to reach for Protobuf, and the Lupyd post makes several of them well.

Deterministic Encoding Isn’t Optional for Crypto

The strongest argument in the entire post is about Firefly’s end-to-end encryption. When you’re computing signatures over serialized data, determinism isn’t a nice-to-have, it’s the entire ballgame.

JSON serialization is famously non-deterministic across platforms. Key ordering varies between implementations. Whitespace handling differs. Floating-point formatting is a crapshoot. If a JavaScript client signs a message and a Rust client verifies it, any discrepancy in serialization breaks the signature.

Protobuf solves this. Fields are ordered by tag number, packed predictably, and produce identical byte sequences across every runtime. For cryptographic workflows, that’s not just useful, it’s necessary.

The same logic applies to their anonymous feedback platform, Noon, which relies on RSA blind signatures. The submission payload must hash identically regardless of whether it originates from a browser, desktop client, or mobile device. You cannot achieve that with JSON. Period.

The Wire Size Savings Are Real

The Lupyd team claims payload reductions of 60% to 80% versus JSON, and that’s consistent with what most teams see. When you’re moving large volumes of data, or you’re on mobile networks where every kilobyte affects battery life and latency, those savings matter.

Cross-Platform Code Generation Prevents Drift

Maintaining data models across TypeScript, Rust, Kotlin, Swift, and Go is a nightmare. A typo in one client repository becomes a production incident nobody can reproduce. Protobuf’s code generation turns those errors into compile-time failures. That’s a real, tangible win.

Where the Binary-First Argument Starts to Crumble

Here’s the problem: the same blog post that makes these legitimate arguments also reveals where the philosophy breaks down.

The “Performance” Argument Has a Blind Spot

One Reddit commenter raised an important point about CPU overhead at data-center speeds, noting that constant encoding/decoding/copying of bytes can eat significant CPU. The Lupyd CEO’s response, that zero-copy libraries exist and varint decoding is “objectively faster than JSON”, is technically true but misses the bigger picture.

At the scale most teams operate, JSON parsing is not the bottleneck. Your database query is slower. Your network round trip is slower. Your TLS handshake is slower. The JSON parser is chewing through microseconds while your API call is burning milliseconds elsewhere. If you’re at Google-scale processing billions of requests per second, sure, optimize the serialization. For everyone else, you’re optimizing a system that isn’t the constraint.

The Debugging Tax Is Real

The Lupyd post acknowledges this: their debugging fallback is content negotiation. If an external service sends Accept: application/json, the backend can serialize the same internal model into JSON. That’s a pragmatic concession, but it also reveals the hidden cost, you’re now maintaining two serialization paths for every endpoint.

Diagram illustrating the use of Protocol Buffers across web, mobile, and database layers
Lupyd’s approach: Protocol Buffers across the entire stack

Meanwhile, your own tooling suffers. You can’t curl an endpoint and read the response. You can’t grep logs for a field value. You can’t inspect a database row without writing a decoder. Every debugging session becomes a trip through a binary decoder, and the “time saved” on serialization gets eaten alive by the time lost on observability.

MessagePack Is the Elephant in the Room

One commenter suggested that MessagePack would be a more developer-friendly alternative for many of these use cases. The counter-argument, that MessagePack lacks schema validation and just pushes type checking onto the developer, is fair, but it raises a question: why is Protobuf the only binary format under consideration?

Because the real value proposition isn’t binary serialization. It’s the schema. The .proto file is the contract. The code generation is the developer experience. The wire format is almost incidental. If the schema is the value, you’re really talking about a contract-first development workflow, and there are lighter tools that deliver that without the binary tax.

The “Banana Republic” Problem

There’s a pattern in software architecture where teams adopt tools for their primary use case, then find increasingly questionable secondary uses. We saw it with microservices, suddenly every internal process needed its own service. We’re seeing it now with event-driven architecture, where Kafka clusters sprout for workflows that could’ve been a cron job.

Protobuf is following the same trajectory. It’s designed for service-to-service communication where performance matters and schema evolution is a real concern. That’s its home turf. But using it for authentication tokens, database storage, and every client-facing HTTP endpoint is the architectural equivalent of using a chainsaw to cut butter.

The weird part is that the Lupyd team stumbled on the best argument against their own approach. When they describe storing Protobuf wire formats in PostgreSQL BYTEA columns, they’re trading away one of the relational database’s core strengths. You can’t query by field value. You can’t index on a nested property. You’re converting the database into a glorified key-value store with extra steps.

The Self-Description Problem

Armin Ronacher’s recent work on deser, his Rust serialization library that deliberately excludes non-self-describing formats, is the sharpest critique of the Protobuf-everywhere philosophy. He notes that Protobuf cannot be supported in Deser because it’s not self-describing, and his entire design philosophy is built around format-driven deserialization.

The practical consequence: you can’t decode a Protobuf message without knowing the schema ahead of time. JSON carries its structure with it. Protobuf doesn’t.

This matters more than most teams anticipate. Your logging infrastructure needs the schema to decode messages. Your debugging tools need the schema. Your data pipeline needs the schema. Your API consumers need the schema. Every consumer becomes coupled to a .proto file version, and schema distribution becomes its own operational problem. It’s a governance overhead that JSON doesn’t impose.

Where the Line Actually Is

So what’s the honest takeaway? Protobuf has a sweet spot, and it’s narrower than its evangelists advertise:

  • Use Protobuf for internal service-to-service calls where both sides control the schema and throughput matters.
  • Use Protobuf for cryptographic workflows where deterministic encoding is a hard requirement.
  • Use Protobuf for high-volume data pipelines where wire size and decode speed genuinely impact infrastructure costs.
  • Client-facing APIs where third-party developers will consume your endpoints. JSON’s self-description is a feature, not a bug.
  • Database storage unless you’re willing to sacrifice queryability and embrace full-table scans.
  • Authentication tokens where the seven lines of base64 decoding you save aren’t worth the debugging friction.
  • Debug logging or observability where human readability is the entire point.

And critically, consider whether you need Protobuf at all. The design constraints of inter-service contracts are often better served by simpler tools. If you’re in a monolith, JSON plus a validation library gives you 90% of the safety with none of the ceremony.

The Actual Enemy Is Indiscriminate Adoption

The worst thing about the Protobuf-everywhere movement isn’t the wire format or the tooling costs. It’s the mindset that a single serialization format should rule every layer of the stack. That’s how teams end up with the kind of architectural debt that takes years to unwind.

Every tool has a purpose. Protobuf’s purpose is efficient, schema-enforced communication between systems that both control their data contracts. When you force it into layers where human readability matters, where dynamic data is the norm, or where observability is a priority, you’re not being “type-safe” or “binary-first”, you’re just making your life harder.

The “Protobuf Everywhere” story is a default choice made by teams who’ve adopted a pattern without asking whether it fits each specific case. The better approach is to treat serialization formats like any other architectural decision: weigh the trade-offs, consider the operational costs, and pick the right tool for the specific job.

Sometimes that’s Protobuf. More often than the evangelists admit, it’s just JSON. And that should be a perfectly acceptable answer.

Share:

Related Articles