The handshake is ordinary HTTP until it is not
A WebSocket connection begins as an HTTP GET with headers that ask the server to upgrade: Upgrade: websocket, Connection: Upgrade, and Sec-WebSocket-Key holding a random base64 value. The client also lists a protocol version and, when it wants one, a Sec-WebSocket-Protocol value naming a transport format.
The server answers with 101 Switching Protocols and a Sec-WebSocket-Accept header computed by hashing the client key with a fixed GUID and encoding the digest in base64. After that response the same TCP connection carries frames instead of HTTP messages. Custom headers on the handshake, such as an authorization token or a device identifier, are the app's authentication, and a client that omits them gets a 401 or an immediate close.
Proxies and load balancers sit between the app and the server on many deployments, and they change the handshake. A proxy may strip custom headers, require a different path, or terminate the connection after a fixed idle period. When a handshake works on one network and fails on another, the proxy is the first thing to check.
Framing and what it means for replay
Every frame starts with a small header. The first byte holds a FIN bit and an opcode, and the second holds a mask bit and a payload length. Lengths below 126 fit in that byte, lengths up to 65,535 use two more bytes, and larger payloads use eight. Client frames are masked with a random four-byte key that is XORed against the payload.
- Opcode 1 is text and opcode 2 is binary.
- Opcode 0 continues a fragmented message.
- Opcodes 8, 9 and 10 are close, ping and pong.
The mask is random per frame and the server does not expect a particular value, so masking does not prevent replay. What prevents replay sits inside the payload: sequence numbers, nonces, timestamps, or a session bound to the connection. Capture tools that show text frames hide the framing, which is convenient until a binary protocol needs frame-level inspection.
Compression is negotiated during the handshake with an extension header, after which payloads may arrive compressed per message. A client that reads the payload without honoring the extension sees binary noise instead of text, and the deflate context is shared across messages, so a single frame cannot be decompressed in isolation.
Finding the message format
Text frames are almost always JSON, and the frames right after the handshake establish the session and name the channels the app subscribes to. Read those before anything else, because they tell you which operations the connection carries.
Binary frames carry a serialized structure. Common choices are protobuf, MessagePack and a hand-rolled layout with a message type in the first bytes. Look for a discriminator: a string field named type, op, cmd or action in text frames, or a leading integer in binary ones. Then find the serializer in the APK. Generated protobuf classes, a MessagePack codec, or a class with pack and unpack methods all point at the layout, and the field numbers or type constants tie the code to what you see in traffic.
Heartbeats matter more than they look. A server that expects a pong within a fixed interval closes the connection when it does not arrive, which produces a failure that resembles an authentication problem. Application-level pings are separate from protocol pings and often carry a payload the server echoes back.
Rebuilding a working client
Start with the handshake and reproduce the headers exactly. Then send the captured frames in order and log at the frame level rather than the message level. A reconnect is a new session, so the client must repeat the authentication message and re-subscribe to every channel it had open. Delivery is ordered within a connection, and most protocols correlate a response to a request with an identifier in the payload rather than by position.
Once the session works, the remaining work is the message catalogue: each type, its fields, which side sends it, and what the server does when a field is missing. That catalogue is what makes the protocol usable from a program rather than from a captured script.
Close codes are part of the protocol and worth recording. A normal close with a standard code differs from an error close, and a server that closes with a policy code is telling you which rule the client broke.
Related work
Reviewed 28 September 2026 · SReverse research desk