Why HTTP, WebSockets and WebRTC?
a simple walk from how people talk to how machines talk
Start with how we talk in real life
Two people talking need three things. First, a medium that carries the voice (air, or a phone line). Second, an address to reach the right person (their phone number). Third, a common language, so the words mean something.
Take a phone call. We dial a number, the call goes through towers and cables, and the friend picks up. The medium and the address are working perfectly. But if we say "kaise ho?" and the friend doesn't know Hindi, nothing happens. The voice reached him, but he can't understand it. If we say "how r u?", he understands and replies "I'm good".
The internet has exactly the same split, and almost everything in this topic comes from it. Some protocols carry the data (medium and address), and some protocols give the data meaning (language).
The carrying part: IP, TCP and UDP
Physical layer. Cables, wifi and 4G/5G towers turn bits into signals (radio waves, light, voltage) and move them. They don't care what the bits mean, just like the tower doesn't care what language we speak.
IP (IPv4 and IPv6). This is the phone number. Every machine gets an address like 142.250.77.46 (IPv4). IPv6 is the longer version, created because IPv4 addresses were running out. IP only delivers on a best-effort basis. A packet can get lost, duplicated, or arrive out of order, and IP won't tell anyone.
Ports. The IP address finds the machine, but one machine runs many apps. A port finds the right app (80 for HTTP, 443 for HTTPS), like an extension number after the main phone number.
TCP and UDP. Both sit on top of IP and use ports. They differ in how they deliver.
- TCP is like a proper phone call. First we confirm the other side is there (the 3-way handshake: SYN, SYN-ACK, ACK). Then every chunk is numbered, the receiver confirms each one (ACK), anything lost is sent again, and everything is handed over in the right order.
- UDP is like shouting a message across the street. No call setup, no confirmation, no re-sending. Some words may not arrive, but it is fast and light.
The important limit is this: TCP gives only a stream of bytes. It has no idea where one message ends and the next begins, or whether the bytes are a name, a photo, or hii. Just like the phone call, the connection is working, but the meaning is missing.
The language part: why HTTP?
When we send hii to a server, the server receives bits. A server is just a program, and it can't "understand" anything by itself. It can only read data in a format it was written to expect, and both sides must agree on that format beforehand. That agreement is a protocol, and for the web it is HTTP.
HTTP is a text format. A request looks like this:
GET /myprofile HTTP/1.1
Host: myapp.com
Authorization: Bearer <token>
The first line has three parts:
-
Method: what we want to do.
GETis "give me something",POSTis "create something" (like sending an email and password to make an account),PUT/PATCHupdate, andDELETEremoves. -
Path: what we want it done to (
/myprofile). - Version: which dialect of HTTP we speak.
The headers add details such as who we are and what formats we accept. The server reads this, knows "ohh, I have to send the profile", runs its code, and replies in the same agreed format:
HTTP/1.1 200 OK
Content-Type: application/json
{"name": "Rahul", "age": 21}
The 200 is a status code that says how it went. 404 means not found, 401 means not logged in, and 500 means the server crashed.
That is the whole answer to "why HTTP": TCP made the call, and HTTP is the language that tells the server what the client wants. This is the same as "how r u" telling the friend what to answer. Two details are worth knowing:
- HTTP is stateless. Each request stands alone, and the server doesn't remember the previous one by itself. That is why cookies, sessions and tokens exist.
- HTTP/1.1 and HTTP/2 run on TCP. HTTP/3 runs on QUIC, which is built on UDP with reliability added on top. So "HTTP means TCP" is the common case, not a law.
- HTTPS is the same HTTP wrapped in TLS, which encrypts it.
HTTP only answers, it never starts the talk
HTTP follows the client-server model with a strict rule: request first, response after. GET /myprofile makes the server send the profile. POST with an email and password makes the server create the account. The server never does anything on its own. It can't say "hey, new message!" unless a request is there to answer.
For loading a profile or checking a balance, this is perfect. But what about a chat? We asked a friend something, and now we want to know whether he has replied. The server won't tell us, so we have to ask.
Polling: asking again and again
Short polling means sending a request every fixed interval, say every 2 seconds, with the question "any new message?". Each time, the server queries the database and replies, and most of the time the answer is "nothing". Why this hurts:
- Each request carries full HTTP headers in both directions, often hundreds of bytes with cookies and tokens.
- Each one makes the server run its code and hit the database.
- The reply can still be up to 2 seconds late, because a message that arrives just after a check waits for the next one.
- At scale it is huge. One million users polling every 2 seconds is 500,000 requests per second, nearly all of them returning nothing.
Throttling is how servers protect themselves from this. The server (or a gateway in front of it) limits how many requests one client may send, for example 100 per minute. If we cross it, we get 429 Too Many Requests. So throttling is a deliberate limit on the client's requests, not the server slowing itself down.
Long polling: let the server wait
Instead of the server answering "nothing" immediately, what if it just holds our request open? The client asks "any new message?". If there isn't one, the server doesn't reply yet. It parks the request and replies only when a message arrives, or when a timeout (typically 20 to 60 seconds) runs out and it sends an empty reply. In both cases, the client sends the next request immediately, so there is nearly always one waiting request at the server.
How does the server wait without wasting effort? It doesn't loop and keep checking the database, since that would only move the waste inside the server. It uses event-driven programming: the handler registers "wake me when a message arrives for this user" and goes to sleep. When a message arrives, an event fires and the server completes the waiting request. This is why event-driven runtimes (Node's event loop, async I/O, epoll) can hold thousands of parked connections cheaply.
This fixes much of the waste and the delay, but it is only a partial fix:
- Every message still costs one full request and response, with all the headers.
- After every reply, the client has to reconnect.
- It is still effectively one-directional. While a request is parked, sending our own message needs a second request.
- Proxies and load balancers may cut long-held connections, and messages arriving between one reply and the next request must be remembered by the server so they aren't missed.
WebSockets: one connection, both directions
The real fix is to stop making a new request for every message. Open one connection and keep it open for the whole chat, so that either side can send whenever it wants. This is called full-duplex communication, and it is what WebSockets give us.
A WebSocket doesn't replace HTTP completely. It starts with HTTP, once, to open the door. The client sends a normal HTTP request asking to switch protocols:
GET /chat HTTP/1.1
Host: myapp.com
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==
Sec-WebSocket-Version: 13
If the server agrees, it answers:
HTTP/1.1 101 Switching Protocols
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Accept: s3pPLMBiTxaQ9kYGzzhZRbK+xOo=
101 Switching Protocols means "agreed". The Sec-WebSocket-Accept value is calculated from the client's key, proving the server really understands WebSocket. Starting as HTTP also helps because the connection passes through the same ports (80 and 443) and proxies that already handle web traffic.
After that exchange, the same TCP connection stops speaking HTTP. Both sides now send small frames (text, binary, ping, pong or close) with a header of only 2 to 14 bytes, instead of hundreds of bytes of HTTP headers on every message.
What this gives us:
- The server can push: "your friend replied" arrives the moment it exists, without any request.
- No repeated headers, no reconnecting, no empty replies.
- URLs are
ws://(plain) andwss://(encrypted with TLS, the one used in production). - Ping/pong frames check the connection is alive, because silent connections get dropped by routers and proxies.
- The trade-off is that the server now holds state for every open connection. Scaling across several servers needs an extra layer (like Redis pub/sub) to deliver a message to whichever server holds the receiver's connection.
- If only the server needs to push (notifications, live scores), Server-Sent Events is a lighter one-way option over plain HTTP.
WebRTC: when TCP's reliability becomes the problem
WebSockets are great for chat because every message must arrive, and in order. That is exactly what TCP guarantees. But live voice and video are different.
Suppose a packet is lost. TCP will not hand later packets to the app until the missing one is re-sent, even if they have already arrived. This is head-of-line blocking. For a download, that is correct. For a live call it is a disaster, because by the time the re-sent packet arrives, that moment of the conversation is 300 ms old and useless. The call doesn't show an error. It freezes and lags, then rushes to catch up.
So for live media we want UDP: if a few milliseconds of voice are lost, skip them and keep going. For a call, fresh and slightly broken is better than perfect and late.
That is why WebRTC exists. It isn't one protocol but a bundle that lets two browsers or apps talk peer-to-peer, directly and not through our server:
- Signaling: before the call, the two sides must exchange who they are, which codecs they support (SDP offer and answer), and how to reach each other (ICE candidates). WebRTC doesn't define how to carry this, so apps usually use a WebSocket or HTTP through a signaling server. That is why WebSockets and WebRTC often appear together.
- NAT traversal: most devices sit behind a router with a private address, so they can't just dial each other. A STUN server tells each peer its public address. If a direct path is impossible (strict firewalls), a TURN server relays the media instead.
- Media: audio and video travel as SRTP (encrypted RTP) over UDP. Arbitrary data (chat, files, game state) travels over data channels using SCTP over DTLS, also on UDP.
- Because UDP gives no help, WebRTC adds its own lightweight control: sequence numbers, a jitter buffer to smooth uneven arrival, congestion control to lower quality on a bad network, and optional re-sends only for important video frames.
Why not WebSockets over UDP?
This question comes up naturally, and the answer has four parts:
- WebSocket is defined on top of TCP. It begins with an HTTP upgrade, and its frame format assumes bytes arrive complete and in order. It can't survive packet loss, so a "WebSocket over UDP" would really be a different protocol.
- Browsers don't allow raw UDP sockets. A page that could send arbitrary UDP packets could be used for attacks and network scanning. WebRTC is the controlled exception, with built-in encryption and a check that the other side agreed to receive.
- Raw UDP gives none of the extras. NAT traversal, encryption, congestion control, jitter buffering and codecs would all have to be built from scratch. WebRTC ships them together.
- They are built for different jobs. Chat, notifications and live dashboards need every message in order, so TCP's guarantees are a feature there.
A newer option is WebTransport, which runs over HTTP/3 (QUIC, built on UDP). It gives browsers both reliable streams and unreliable datagrams, a bit like a "WebSocket that can also skip reliability".
The whole story in one list
- Load a profile, create an account: use HTTP. One request, one response. Only the client can start.
- Check for news on a timer: use short polling. Repeated requests, mostly empty replies. Only the client can start.
- Check for news, with less waste: use long polling. The request is parked until data arrives. Only the client can start (the server replies late).
- Live chat, notifications, live scores: use WebSocket. One TCP connection kept open, data sent as frames. Both sides can start.
- Live voice, video, games: use WebRTC. Peer-to-peer, mostly over UDP. Both sides can start.
In one line each
- IP, TCP and UDP are the phone line: they carry the bits.
- HTTP is the language: it gives the bits a meaning (
GET /myprofile= send my profile). - Polling and long polling are workarounds for HTTP only answering when asked.
- WebSocket keeps the line open so both sides can speak anytime.
- WebRTC gives up TCP's guarantees on purpose, because in a live call, a late word is worth less than a missing one.
Top comments (0)