DEV Community

Cover image for When a VPN Says "Connected" but Traffic Is Dead: Building Adaptive Connect 2.0
Colitu VPN
Colitu VPN

Posted on Originally published at colitu.com

When a VPN Says "Connected" but Traffic Is Dead: Building Adaptive Connect 2.0

A connection is not proof of connectivity. Here's what we learned while building a VPN that has to survive networks that silently stop forwarding traffic.

A VPN can establish a tunnel, complete its handshake, display a reassuring green Connected indicator, and still fail at its most important job: carrying traffic.

This is not a theoretical edge case.

While developing Colitu, an open-source VPN project focused on reliable connectivity in restrictive network environments, we encountered a frustrating failure mode: some transports connected successfully, then approximately 30–50 seconds later, application traffic stopped flowing.

The TCP connection remained open. The VPN appeared connected. But websites would not load.

Restarting the connection sometimes restored traffic, only for the problem to return.

A successful handshake wasn't enough. A low-latency server wasn't enough. Even a backup route wasn't enough if the health monitor misunderstood what was happening.

These experiences became the foundation of Colitu Adaptive Connect 2.0.

In this article, we'll cover the engineering decisions, the mistakes, the experiments, and the problems we haven't solved yet.


1. The problem: a tunnel that looks alive but isn't

Many VPN clients follow a familiar process:

  1. Measure server latency.
  2. Select a server.
  3. Establish a tunnel.
  4. Report that the VPN is connected.
  5. Reconnect when the tunnel terminates.

That strategy assumes transport failure will be obvious. In restrictive networks, it often isn't.

On one network, we observed VLESS Reality, VLESS XHTTP, and Trojan connections successfully establish sessions and then stop carrying traffic after roughly 30–50 seconds. Hysteria2 continued working on the same network.

The important distinction is between transport state and application-level reachability.

A socket remaining open does not prove that useful traffic is passing through it.

This exposed a weakness in Adaptive Connect 1.x: in-session monitoring covered Hysteria2 but not every transport. We needed a health model based on what users actually depend on.

2. Verify real traffic, not just handshakes

Adaptive Connect 2.0 treats a connection as functional only after useful traffic successfully passes through the selected tunnel.

Instead of relying exclusively on Colitu infrastructure, connection-time checks use established external connectivity endpoints to help distinguish three states:

State Why it matters
Transport established The protocol completed a connection attempt.
Application traffic verified Actual requests can traverse the VPN.
Underlying network unavailable The device itself may have lost access.

Each attempt has a bounded time budget. Failed candidates are followed by alternatives rather than endless retries. On Android, the UI presents a verification phase instead of immediately claiming success.

A critical detail: the primary route has its own verification path.

If normal traffic is already passing through a backup, a generic health check might report success even while the primary route is broken. The dedicated local-only verification entry point tests the primary route without letting the backup conceal a failure.

That distinction informs future transport decisions.

3. Give each network its own memory

A transport that fails on home Wi-Fi may work perfectly over mobile data. Adaptive Connect 2.0 keeps bounded, device-local history based on network context.

Observation Retention
Last working server 24 hours
Last working transport per server 24 hours
Transport freeze marker Normally up to 6 hours
Server-wide failure 30 minutes

The local history is capped at 200 entries.

These limits are intentional. A failed connection should affect subsequent ranking, but it should not make a transport permanently unusable.

Failure history should guide exploration, not make recovery impossible.

4. Learn from network conditions without collecting browsing history

Local memory helps a device make better decisions. Aggregated network observations can also help identify where particular transports are failing.

Adaptive Connect 2.0's reporting design uses:

  • HMAC-signed network tokens carrying country and ASN information, valid for 48 hours.
  • Hourly success/failure counters grouped by country, ASN, and transport.
  • Temporary keyed pseudonyms to count distinct devices.
  • Seven-day retention of aggregated counters.

This design does not require URLs visited, browsing history, or application payloads to produce transport recommendations.

For example, the ASN-level rule requires observations from at least five devices and a success rate below 20% over the previous 24 hours before deprioritizing a transport. With insufficient evidence, the system falls back to broader signals rather than declaring a protocol blocked.

These are engineering safeguards, not a claim of perfect anonymity or immunity to malicious telemetry.

5. Our first backup strategy was wrong

At first, transport diversity seemed like the natural basis for failover.

If the primary uses UDP, choose a TCP-based backup:

Primary: Hysteria2 (Server A)
Backup:  Trojan   (Server B)
Enter fullscreen mode Exit fullscreen mode

That sounds sensible: different transports might fail differently.

But on the network we were testing, multiple TCP-based transports suffered the same freezing behavior. A different protocol was not necessarily a working alternative.

We had confused protocol diversity with failure independence.

The revised strategy prefers a transport already proven to work on the current network, even when that means using the same transport on a different server:

Primary: Hysteria2 (Server A)
Backup:  Hysteria2 (Server B)
Enter fullscreen mode Exit fullscreen mode

This was not about declaring Hysteria2 universally superior. It was about choosing the best observed independent route for the network at hand.

Lesson: Protocol diversity does not guarantee failure independence.

6. A backup-aware health monitor

One of our early failure-injection tests exposed a subtle bug.

We disabled the primary route. After approximately 12 seconds, the backup began carrying traffic. That looked like successful failover.

But the old health monitor still saw the failed primary, classified the whole VPN as unhealthy, and restarted the tunnel.

The watchdog destroyed the recovery the backup had achieved.

We changed the health model to track two separate questions:

  1. Primary health: Can the original route carry traffic?
  2. Effective connection health: Can traffic still pass through any usable route, including the backup?
Primary route Effective route Correct response
Healthy Healthy Continue normally.
Failed Healthy Preserve the session; record primary failure.
Failed Failed Start recovery or reconnection.

The backup is configured inside the running tunnel and used on demand. It is not a continuously active second VPN tunnel.

We also encountered an Xray observer/balancer detail: prefix-based tag matching could accidentally include the backup route in a primary-only health check. Making primary and backup tags non-overlapping removed that ambiguity.

This is why a failover design cannot be validated just by reading its configuration.

7. Parallel probing without unbounded retries

Sequentially trying every transport on every server can be painfully slow when packets are silently dropped.

Adaptive Connect 2.0 uses bounded selection:

  • The best two connection candidates may be tested in parallel.
  • The first candidate that successfully carries real traffic wins.
  • Automatic server fallback is limited to three servers in one connection attempt.
  • Transport, server, and overall attempts each have time budgets.
  • Previous network observations influence the initial ordering.

The parallelism is inspired by the Happy Eyeballs approach described in RFC 8305.

This is not a newly invented connection algorithm. The engineering value is in combining bounded candidate racing, network-specific history, and real-traffic verification.

8. Testing the system by deliberately breaking it

We built an internal fault-injection tool to selectively disrupt one client's access to a server or transport, without intentionally affecting other users.

The experiment took place on October 9–10, 2026, using one Android 11 phone connected to one home Wi-Fi network in an environment where DPI was present.

In the 60-minute endurance test, we used eight servers and injected 11 controlled failures lasting roughly 60–120 seconds each. An HTTP probe was sent through the VPN approximately every five seconds.

Results

Measurement Result
Test duration 60 minutes
Servers involved 8
Controlled fault injections 11
Successful HTTP probes 741 / 750
HTTP probe success rate 98.8%
Primary-server outage 101 seconds
Interruption observed during that outage ~5 seconds
Longest observed interruption 40 seconds

The most encouraging event was a 101-second outage of the active primary server. The VPN remained in its connected state, the backup carried traffic, and the observed interruption was approximately five seconds.

However, a separate test revealed a hard limit: when Hysteria2 was disabled across all tested servers, the same network's TCP transport freezes left no usable alternative. That caused a 40-second interruption.

Some injected faults targeted servers or transports that were not carrying the active session. Those produced zero observed downtime, but should not be represented as successful active-path recoveries.

We believe that distinction matters more than a single flattering availability percentage.

Important test limitations

This was:

  • One device and one home Wi-Fi network.
  • One hour of controlled experiments.
  • Mostly injected server/transport failures, not a comprehensive reproduction of real-world censorship.
  • Not a cross-platform or independently verified reliability benchmark.

The 98.8% figure describes successful HTTP probes within this specific test, not worldwide Colitu availability.

9. What Adaptive Connect 2.0 does not solve

Existing TCP sessions are not guaranteed to survive server failover.

Keeping the VPN interface active and keeping an established application connection alive are different engineering problems. A download may still stop when its underlying exit path changes.

We have not proven the same results across every platform or network.

Different mobile operators, Wi-Fi environments, devices, and censorship systems may produce very different outcomes.

Controlled fault injection is not equivalent to every real censorship event.

Actual interference can be less predictable and may respond dynamically to a client's behavior.

No software can route traffic through a path that does not exist.

If every accessible route is unusable, Adaptive Connect cannot manufacture connectivity.

10. Next: session continuity across transports

Our next research question is harder than reconnecting a VPN:

Can existing application sessions survive a transport change without having to reconnect?

The proposed Colitu Session Layer (CSL) separates a logical session from its underlying transport. Early laboratory work is promising, but CSL is not a production capability of Adaptive Connect 2.0.

There is also an important distinction between switching protocols on one server and moving an established session to another server with a different egress address. The latter requires additional session and egress management.

We are approaching that work as research, not as a feature we can already promise.

What we learned

The biggest lesson from Adaptive Connect 2.0 is not that one VPN protocol is better than another.

It's that connectivity is a changing, measurable property of a network.

  • A working socket is not proof of working traffic.
  • A different protocol is not necessarily an independent backup.
  • A health monitor that ignores backup behavior can trigger avoidable outages.
  • Failure observations should inform decisions without preventing exploration.
  • Controlled experiments should report failures as clearly as successes.

Reliable connectivity is less about choosing the perfect tunnel and more about continuously verifying that a usable path still exists.

That's the problem we're trying to solve at Colitu.


Read the research and source code

We're interested in feedback from developers working on networking, transport protocols, privacy-preserving telemetry, fault injection, and censorship circumvention.

What failure modes would you test next? How would you design a health monitor that distinguishes a dead primary from a healthy session on its backup?

Colitu Engineering

Built for networks that fight back.

Top comments (2)

Collapse
 
m_montazeri profile image
Mohammad Montazeri •

One next test would be a verification-endpoint outage while the tunnel still carries useful application traffic. In a controlled fixture, block or delay only the probe destination, keep an independent application request and a long-lived connection running, and record whether the monitor restarts a working session. That exercises the opposite error from a dead tunnel reporting healthy.

I would also alternate probe success/failure near the timeout boundary and record the state-transition timeline: detection, backup selection, recovery and any switch back. That should expose flapping and show whether the failure/recovery thresholds are different enough to avoid repeatedly discarding a usable backup. Report the injected condition and the application interruption separately from the probe success percentage.

These are proposed tests, not results from running Colitu. AI-assisted technical feedback from Lisar, a VPN service.

Collapse
 
colitu profile image
Colitu VPN •

Thanks, both are good tests and we'll run them. Here is the plan.

Probe-endpoint outage with a working tunnel. In a test build we'll point verification at an endpoint we control, then make it time out or slow down while the tunnel and everything else stay healthy. During the outage we'll keep independent application requests and a long-lived download running, and record whether the monitor restarts or switches a session that is still carrying traffic.

Flapping near the timeout boundary. The same endpoint will alternate success and failure just below and just above the probe timeout, in a few patterns. We'll log the full state timeline: detection, backup selection, recovery and any switch back. That should show whether the failure and recovery thresholds are far enough apart to keep a usable backup.

As you suggested, we'll report the injected condition and the application-level interruption separately from the probe success rate. If the monitor needs changes, we'll make them and re-run before publishing. The results will go up on Colitu Lab next to the first report: lab.colitu.com/research

Thanks again to you and the Lisar team for the constructive feedback.