SReverseby Simpa Labs

Binary protocols

Reverse engineering protobuf APIs on Android

A protobuf body is readable without its schema, but only as numbered fields. Reconstructing the .proto turns those numbers into names your code can use.

The schema is not in the payload

A protobuf message carries field numbers and values. The names, the types and the nesting live in a .proto file that both sides compiled into their own code. When you capture the bytes you hold the values and the numbers, and the names are the part you have to recover.

Every field is a tag followed by a value. The tag is a varint whose low three bits are the wire type and whose remaining bits are the field number: (field_number << 3) | wire_type. Wire type 0 is a varint, 1 is 64-bit, 2 is length-delimited, and 5 is 32-bit. Wire types 3 and 4 are the legacy group markers, which current code rarely emits.

Reading a message with no schema

You can decode a protobuf body with nothing but the wire format. Walk the bytes tag by tag, read each value according to its wire type, and print the number and the raw value. The result is a tree of numbered fields rather than names, and it is a complete view of what the app sent.

  • A varint that decodes to a small integer is usually a field number you can match against a generated constant later.
  • A length-delimited blob that itself parses as tag and value pairs is a nested message.
  • A length-delimited blob that does not parse is likely a string, a bytes field, or an embedded serialized object.
  • A repeated field appears as the same field number more than once, unless the build uses packed encoding for numeric types.

Two details trip people up. A sint32 or sint64 field uses zigzag encoding, so the decoded varint is not the value. And proto3 omits fields set to their default, which makes an absent field and a field containing zero indistinguishable on the wire.

Finding the schema inside the APK

Most Android apps ship the compiled form of their proto definitions. The generated Java classes are the easiest place to start because the field numbers survive as named constants.

  • Look for classes that extend GeneratedMessageV3 or implement MessageLite.
  • Search the DEX string pool for FIELD_NUMBER, getDefaultInstance, parseFrom and newBuilder.
  • Constants such as USER_ID_FIELD_NUMBER = 3 pair a name with a number even when the class body is obfuscated.
  • Enum types carry their own constants, and an unrecognised enum value in a response often prints as a number.
  • Descriptor blobs and file descriptor names sometimes survive as string constants, which recover the package and message names directly.

Protobuf-lite builds strip descriptors but keep the field constants. Kotlin code generated from a proto keeps the same names. If the app uses a different runtime, such as a hand-written codec or a generated wire adapter, the constants may not exist and you fall back to wire-level decoding and matching against live traffic.

Implementing the wire walk takes little code in any language, and the rule that matters is that the cursor must consume every byte. If it stops before the end of the buffer or runs past it, you have misread a length. Length-delimited fields cause most of those errors, because a length is itself a varint and skipping a tag shifts every field after it.

A search for the message class names is often more productive than a search for network code. Proto-generated classes usually sit together in the DEX, so finding one points at the package that holds the rest.

Rebuilding the .proto

Once you have numbers, names and observed values, write a .proto file and compile it. Round-trip the captured bytes through your generated classes: encode the same values and compare the output byte for byte. A byte-identical encoding confirms the field numbers and the types. If the encoder produces different bytes, check field ordering, packed against unpacked repeated fields, and whether a field you marked as a string is actually bytes.

Most runtimes preserve unknown fields, which gives you a verification trick. Feed a captured message through your parser, re-encode it, and see whether the bytes you could not name still survive.

Where protobuf hides in the transport

The bytes rarely travel alone. A request may be an HTTP POST with application/x-protobuf as the content type, a base64 string inside a JSON envelope, a gRPC frame with a five-byte prefix, or one field inside a WebSocket message. Identify the envelope before you decode the message, so you know where the protobuf body starts and ends.

When an endpoint returns protobuf, capture a success and a failure. Error responses often omit optional fields, which exposes the structure with fewer values in the way.

Reviewed 28 September 2026 · SReverse research desk

Start your full APK reconstruction

Send the full APK

Projects start at $120. Choose WhatsApp or email, then attach the APK in the app that opens. We reply within one hour with the next step and send the fixed quote after review.

Want us to contact you?