Size and speed comparison

A microcontroller has a few hundred kilobytes of flash, a few tens of kilobytes of RAM, and a clock in the tens of megahertz. A serialization library has to fit in all three. This page shows what Embedded Proto costs on a Cortex-M4 next to nanopb, the other protobuf library for microcontrollers, and next to three alternatives people consider instead of protobuf: FlatBuffers, CBOR, and JSON. Every number below was measured on the board, with the same messages and the same values in every library.

In short: Embedded Proto is the fastest protobuf library on every message, the fastest of the five libraries on all but the batch, where FlatBuffers wins by reading in place, and the smallest of the protobuf libraries. It serializes and parses a telemetry message in a third of the cycles nanopb needs, and it takes 2.5 to 5.3 kB of flash where nanopb takes 5.6 to 6.9 kB. Its message objects are a few bytes larger than nanopb's structs, but with the buffers and the stack counted the RAM the two need is within a few hundred bytes of each other, Embedded Proto lower in five of the six scenarios. The bytes on the wire are identical, because both speak the same protobuf. The one place nanopb is smaller is a schema with many message types, which the last chapter measures.

Flash [b]Cycles [-]
ScenarioEmbedded ProtonanopbEmbedded Protonanopb
Telemetry, serialize and parse25805640+119%458713787+201%
Telemetry batch, 20 samples38885840+50%84091220239+162%
Configuration with oneofs53006872+30%2563269448+171%
Command and acknowledgement26125648+116%27438818+221%
Firmware chunk, 256 bytes24805904+138%1198416732+40%
Firmware chunk, streamed28245732+103%1778318152+2%

Flash is in bytes above an empty firmware, cycles are for one serialize and one parse of the message at 84 MHz. 4587 cycles is 55 microseconds. The percentage next to a number is its difference to Embedded Proto, in every table on this page.

How the numbers are measured

Every scenario is one C++ function per library. It fills a message with fixed values, serializes it into a byte array, copies the array as if it were received, parses it into a fresh message, and checks every field against the values it started with. A scenario which fails any check reports zero bytes and is not stored. The Embedded Proto and nanopb versions produce byte for byte the same wire data, the others produce what their format produces.

Five things are measured. Flash is the text size of the firmware minus the text size of a firmware which only blinks the LED, so the number is the library plus the generated code plus the scenario, nothing else. Cycles are read from the Cortex-M4 cycle counter around the scenario function, so they include filling and checking the message, not only the library calls. Message is the size of the message object your firmware declares: sizeof for Embedded Proto, nanopb and zcbor, the document plus its heap for ArduinoJson. FlatBuffers has no message object, a FlatBuffer is read in place, so its column is empty. Peak RAM is the most RAM the scenario holds at any moment: the message objects, the buffers, the heap of a builder or document, and the stack the library uses along the way. The firmware paints the free RAM with a pattern before the scenario runs and afterwards finds the deepest word the pattern is gone from, plus how far the heap grew. Wire is the length of the serialized message.

All code is compiled with -Os, -ffunction-sections -fdata-sections and linked with --gc-sections. All C++ is compiled with -fno-exceptions -fno-rtti -fno-threadsafe-statics -fno-use-cxa-atexit, the flags STM32CubeIDE and every other embedded IDE set for a C++ project. Each library runs in the configuration its own documentation recommends for a microcontroller: nanopb with PB_NO_ERRMSG, flatcc with its emitter page sized to the message, as its documentation advises for constrained devices, zcbor with its default encoding, ArduinoJson 7 with its slot pool sized to the document instead of the default 128 slots, a setting its documentation describes for small targets. Embedded Proto runs with its defaults. Sizing the flatcc page and the ArduinoJson pool to the message is the same thing every scenario does with its buffers, and no library was tuned beyond that.

Telemetry

The small message every device sends: seven fields of mixed width, an enumeration, and a negative number. Please note the sint32 for the negative value, which is what protobuf advises.

enum Status
{
  OK      = 0;
  WARNING = 1;
  ERROR   = 2;
}

message Telemetry
{
  uint32 device_id   = 1;
  uint64 timestamp   = 2;
  float  temperature = 3;
  float  humidity    = 4;
  uint32 battery_mv  = 5;
  Status status      = 6;
  sint32 rssi        = 7;
}
LibraryFlash [b]Cycles [-]Message [b]Peak RAM [b]Wire [b]
Embedded Proto258045874838428
nanopb5640+119%13787+201%40−17%528+38%280%
zcbor3136+22%9712+112%64+33%652+70%37+32%
FlatBuffers5528+114%8770+91%—1940+405%60+114%
ArduinoJson12056+367%58917+1184%192+300%1840+379%120+329%

Embedded Proto parses this message in 4587 cycles, nanopb in 13787. The difference is the generated code: Embedded Proto writes a serialize and a parse function per message with the field types known at compile time, nanopb walks a descriptor table with one generic loop. Serializing alone, without the parse, takes 1912 cycles in Embedded Proto and 5891 in nanopb, and 2160 against 2628 bytes of flash. With the two buffers and the stack counted, the round trip peaks at 384 bytes of RAM in Embedded Proto and 528 in nanopb. FlatBuffers needs 1.9 kB for the same 60 bytes on the wire, most of it the builder it constructs a message with.

Telemetry batch

A device which sleeps between measurements sends them in one go. The batch holds twenty samples as nested messages, so this is the scenario for repeated fields and nesting.

message Sample
{
  uint32 offset_ms = 1;
  float  x         = 2;
  float  y         = 3;
  float  z         = 4;
}

message TelemetryBatch
{
  uint32 device_id       = 1;
  uint64 timestamp       = 2;
  repeated Sample samples = 3;
}

The repeated field holds at most twenty samples in every library. In Embedded Proto that is the maxLength option, in nanopb max_count, and the schemas of the other libraries say the same.

LibraryFlash [b]Cycles [-]Message [b]Peak RAM [b]Wire [b]
Embedded Proto3888840915922448408
nanopb5840+50%220239+162%344−42%2236−9%4080%
zcbor3360−14%112420+34%672+14%3020+23%494+21%
FlatBuffers6580+69%51769−38%—4748+94%540+32%
ArduinoJson11948+207%543730+547%1600+170%9556+290%994+144%

Here FlatBuffers is the fastest. It does not parse at all, a FlatBuffer is read in place, so a message which is mostly read wins there. It pays with 4.7 kB of RAM at its peak, twice what the protobuf libraries need, and a third more bytes on the wire. Embedded Proto's message is larger than nanopb's because every nested message carries its own parse state, 12 bytes per sample, which is also why its peak RAM is 200 bytes above nanopb's here. zcbor is the smallest in flash and the closest in speed.

Configuration

The complex message: a oneof per peripheral, nested messages in a repeated field, a string, and an optional scalar with presence. Four peripherals are set, two UARTs, an SPI, and a motor.

message UartConfig
{
  uint32 baud      = 1;
  uint32 data_bits = 2;
  bool   parity    = 3;
}

message SpiConfig
{
  uint32 prescaler = 1;
  bool   cpol      = 2;
  bool   cpha      = 3;
}

message MotorConfig
{
  uint32 max_speed = 1;
  bool   invert    = 2;
  uint32 home_pin  = 3;
}

message Peripheral
{
  uint32 index = 1;
  oneof kind
  {
    UartConfig  uart  = 10;
    SpiConfig   spi   = 11;
    MotorConfig motor = 12;
  }
}

message Configuration
{
  uint32 version                    = 1;
  string device_name                = 2;
  repeated Peripheral peripherals   = 3;
  optional uint32 report_interval_ms = 4;
}
LibraryFlash [b]Cycles [-]Message [b]Peak RAM [b]Wire [b]
Embedded Proto530025632248133667
nanopb6872+30%69448+171%128−48%1456+9%670%
zcbor4856−8%33310+30%164−34%1660+24%88+31%
FlatBuffers8552+61%28965+13%—3444+158%252+276%
ArduinoJson12912+144%148243+478%535+116%3676+175%337+403%

Six message types are in play here, and it shows in the flash of the two generated-code libraries, Embedded Proto and zcbor, which sit close together. Embedded Proto is still the fastest by a margin: a oneof costs it a switch statement, nanopb a search through the descriptor. FlatBuffers has to wrap the oneof in a union of tables, which is why its wire bytes are nearly four times the protobuf ones.

Command and acknowledgement

A tiny pair, a command with a oneof action going down and an acknowledgement coming back. Each is a few bytes on the wire, so the fixed cost of every call shows. The scenario does both round trips.

enum Result
{
  UNKNOWN  = 0;
  ACCEPTED = 1;
  REJECTED = 2;
  BUSY     = 3;
}

message DeviceCommand
{
  uint32 sequence = 1;
  oneof action
  {
    bool   set_output      = 2;
    sint32 move_to         = 3;
    uint32 reboot_delay_ms = 4;
  }
}

message DeviceAck
{
  uint32 sequence = 1;
  Result result   = 2;
}
LibraryFlash [b]Cycles [-]Message [b]Peak RAM [b]Wire [b]
Embedded Proto26122743242729
nanopb5648+116%8818+221%12−50%400+47%90%
zcbor3132+20%7338+168%20−17%532+96%16+78%
FlatBuffers6332+142%14617+433%—1940+613%68+656%
ArduinoJson10720+310%26906+881%265+1004%1008+271%57+533%

Two messages of nine bytes together take Embedded Proto 2743 cycles, 33 microseconds, for two serializations and two parses. FlatBuffers spends most of its 14617 cycles setting up its builder, and 1.9 kB of RAM on two messages of nine bytes, a fixed cost a small message can not amortize.

Firmware chunk

The payload: one chunk of a firmware image, 256 bytes, with its position and a checksum. It is measured twice. First with the bytes resident in the message, the way every library can do it.

message FirmwareChunk
{
  uint32 offset     = 1;
  uint32 total_size = 2;
  bytes  data       = 3;
  uint32 crc32      = 4;
}
LibraryFlash [b]Cycles [-]Message [b]Peak RAM [b]Wire [b]
Embedded Proto2480119842881760273
nanopb5904+138%16732+40%272−6%1876+7%2730%
zcbor3004+21%13534+13%36−88%17640%280+3%
FlatBuffers5868+137%15217+27%—3832+118%296+8%
ArduinoJson10884+339%2564438+21299%2240+678%11396+548%978+258%

zcbor's 36 bytes of message are honest and misleading at the same time: it does not copy the bytes, the message only points at them, so the 256 bytes live somewhere else in your program. The peak RAM column shows it: 1764 bytes, the same as Embedded Proto's 1760. JSON has no bytes type. The chunk goes as an array of 256 numbers, which is why ArduinoJson needs two million cycles and a kilobyte of wire for it.

Then the same chunk streamed through callbacks in windows of 32 bytes, so that the message never holds the payload. Only the two protobuf libraries offer this.

LibraryFlash [b]Cycles [-]Message [b]Peak RAM [b]Wire [b]
Embedded Proto282417783601376273
nanopb5732+103%18152+2%20−67%1424+3%2730%

The message shrinks from 288 bytes to 60, the peak RAM from 1760 to 1376, the bytes on the wire stay the same, and the cycles are about two percent under nanopb's. Most of the RAM left is the two 512-byte buffers the scenario sends and receives through. This is how a device with a few kilobytes of RAM receives a firmware image of any size.

Schema growth

The scenarios above use one to six message types. A real product has more, and each type costs flash. To measure how much, a schema defines twenty message types of the same shape, and three scenarios round trip the first one, the first five, and all twenty of them:

message Growth_07
{
  uint32 id      = 1;
  sint32 value   = 2;
  float  measure = 3;
  bool   flag    = 4;
  string label   = 5;
}
Library1 type [b]5 types [b]20 types [b]Per type [b]
Embedded Proto2892578414592616
nanopb6424+122%7500+30%11380−22%261−58%
zcbor3660+27%6252+8%16196+11%660+7%
FlatBuffers7100+146%8252+43%11908−18%253−59%

Flash in bytes, the last column is the difference between twenty types and one, divided by nineteen. Embedded Proto starts 3.5 kB below nanopb and grows 355 bytes per message type faster, so the two lines cross at about ten message types. Below that Embedded Proto is the smaller library, above that nanopb is. Please note that a type only costs flash when your firmware serializes or parses it. Declaring it in the *.proto file is free, the linker removes the generated code of a message you never use.

The cycles tell the other half: twenty round trips take Embedded Proto 77272 cycles and nanopb 190419. The 616 bytes buy a serialize and a parse function per message, and that is where the speed comes from.

The other formats

FlatBuffers, CBOR, and JSON are not protobuf, so the comparison with them is a comparison of formats as much as of libraries. FlatBuffers reads without parsing and wins where a message is mostly read, at the cost of a builder which takes 1.5 to 2.3 kB of RAM more than Embedded Proto in every scenario, and larger messages. zcbor, the CBOR library from Zephyr, is the closest to Embedded Proto in design and in results: a little larger on the small messages and a little smaller on the batch and the configuration, slower on everything, and more bytes on the wire, a third more on the telemetry message. ArduinoJson is the JSON library everyone starts with, and the table shows what a text format costs: two to six times the bytes, six to thirteen times the cycles, over two hundred times on the firmware chunk, and one to eleven kilobytes of RAM at its peak.

Getting the same results in your project

Three things make these numbers, and all three are in your hands.

First, the compiler flags. Compile your C++ with -fno-exceptions -fno-rtti and let the linker remove unused code with -ffunction-sections -fdata-sections and --gc-sections. Without the first two, GCC links its exception unwinder and type information into any C++ firmware, which costs about 4.5 kB whether or not the code throws. Every embedded IDE sets these flags for a new project, a CMake or Makefile project has to set them itself. The page on code size has the details.

Second, the sizes of your fields. Every repeated, string, and bytes field has a maximum length, given as a template parameter or in the options file. That length is what the message object takes in RAM, no more. A payload larger than your RAM goes through callback storage instead, as the firmware chunk above does.

Third, the number of message types your firmware uses. Each one costs about 600 bytes of flash. With a handful of messages Embedded Proto is the smallest protobuf library you can pick, with dozens nanopb is smaller, and in both cases Embedded Proto is the faster one. Take the growth table, count your messages, and you know where you land.

Do you want to see the numbers for your own messages? Clone the comparison, put your *.proto file in the Protofiles folder, write a scenario for it, and run the tool. The README of the project describes how.