SReverseby Simpa Labs

Binary protocols

Recovering field boundaries in a custom binary protocol

When an app speaks a binary protocol of its own design, no schema ships with it. The format has to be inferred from the bytes, and the reliable way to infer it is to change one byte at a time and watch what happens.

Recover the framing before the meaning

A binary message is a sequence of bytes whose structure the client defines. Recovering that structure means answering where a message starts and ends, how a field declares its own length, which type a value has, and where one field stops and the next begins. Field names and semantics come later, and often they never matter, because a client that reproduces a request only has to build the same bytes.

Work from two captures of the same operation with one obvious difference in values. The shared prefix is framing, the bytes that differ are data, and the position where the difference starts is a boundary candidate. A tool that displays hex with offsets and an ASCII column is enough for this pass.

Length prefixes draw the first map

A length prefix states how many bytes belong to a field or to the whole message, and it is the most common way a protocol makes itself parseable. A prefix can be one byte, two bytes or four, and it can be little endian or big endian. A single byte capped at 255 is common for strings, and a four byte signed length is common for whole payloads.

The test is direct. Read the candidate as an integer in each width and byte order, then check whether that many bytes after the prefix land on a plausible boundary. A value that equals the distance to the end of the message, or to the start of the next repeated structure, is a length rather than a coincidence.

Tag and type byte pairs

Protocols that model messages as a set of optional fields emit a tag identifying the field and a type byte telling the reader how to interpret what follows. The type decides the length, so the type table comes before any parsing.

Type bytes usually fall into a small pattern. One low value may mean a variable length integer, another a fixed width integer, and another a length delimited block of bytes. Once you have two captures, you can match type bytes with the widths that follow them and build the table. Tag numbers then appear as small integers that repeat in the same order across captures of the same call.

Nested messages are tagged fields whose payload is itself a message, and their length prefix covers the inner bytes. Recursion into that payload is what turns a flat hex dump into a tree, and it is the step where a protocol description starts to resemble a schema.

Varints, fixed width fields and endianness

A varint stores an integer in as few bytes as its magnitude allows, with the high bit of each byte marking continuation. Small numbers take one byte, and the encoding groups seven bits per byte.

Varints are easy to confirm because the encoded width follows the value. Send two requests that differ only in a numeric field, one small and one large, and the message grows at that position. A field that shrinks when the value shrinks is a varint.

Fixed width fields are the opposite. A four byte value takes four bytes for a value of one and for a value of a million, so the width cannot be inferred from a single capture. Repetition in a fixed width structure leaves regular gaps between recognisable values, and that regularity is the evidence. Endianness stays a coin flip until one field decodes into a date, a length that matches the message, or an identifier you already know.

Alignment and padding

Some formats align fields to powers of two and insert padding to reach the next boundary. Padding shows up as a run of zero bytes between two fields whose values you can otherwise locate, and it moves when the preceding field changes size.

Alignment exists because the producer writes a native structure rather than a packed one, and the padding bytes are not part of any field. Treating padding as data is a common cause of a client that sends a correct body with extra bytes in the middle. Check whether padding follows the value that precedes it, and whether the message length is a multiple of the alignment.

Confirm a boundary by mutating bytes

Inference produces candidates. Mutation confirms them. Change one byte inside the suspected field and send the request, then watch what the server does and what the app does with the response.

  • Change a byte in the middle of a length prefix and the parse fails. A rejected message, a truncated parse or a crash at a fixed offset tells you a length lives there, because the reader trusted the value and walked out of the buffer.
  • Change a byte in the middle of a payload and only that value changes. The neighbouring fields keep their meaning, which places both boundaries outside the byte you touched.
  • Change the first byte of a header block and the message keeps parsing. A field that no integrity check covers can often be changed, which tells you what the check protects.
  • Change a value an integrity check covers and the request is refused. The refusal confirms that the region you touched is signed, which is the information a client needs.

From bytes to a working client

A client that reproduces a binary protocol needs a parser that preserves unknown bytes and an encoder that emits them unchanged, so a field you never identified survives a round trip. SReverse recovers custom binary protocols from the client and delivers a codec with the identified fields typed and the rest kept as raw bytes. When the framing turns out to follow a published scheme, recovering the map of a protobuf message applies the same method against a known wire type.

Reviewed 28 September 2026 · SReverse research desk

Start your full APK reconstruction

Send the full APK

Projects start at $120. Choose WhatsApp or email, then attach the APK in the app that opens. We reply within one hour with the next step and send the fixed quote after review.

Want us to contact you?