An engineering deep dive into fault injection, silent VPN failures, protocol recovery, warm spare reliability, and the uncomfortable difference between implementing a fix and proving it works.
By Colitu Engineering | October 10, 2026
A VPN Can Be Connected and Still Be Broken
There is a particularly frustrating kind of network failure.
Your VPN application says Connected.
The tunnel is established. The process is running. The connection has not been explicitly terminated.
But websites stop loading. Messages remain undelivered. Requests time out.
From the operating system's perspective, the connection may still exist. From the user's perspective, the internet is gone.
This is one of the problems we are trying to solve with Colitu Adaptive Connect 2.0.
Modern restrictive networks do not always block connections outright. Some allow the initial handshake, exchange a small amount of data, and then silently stop forwarding traffic.
Others selectively interfere with UDP, terminate particular TCP connections, introduce packet loss, or make an otherwise healthy server temporarily unreachable.
Supporting multiple protocols is useful, but it does not automatically solve these problems.
A fallback system can select the wrong alternative. A health monitor can mistake an established socket for a working connection. A backup tunnel can silently die without being replaced.
And sometimes the mechanisms designed to improve reliability become the reason recovery takes longer.
We wanted to understand these failures at the level where they actually happen.
So we built a fault-injection experiment around a real Android device on a restrictive residential network in Russia.
We deliberately broke its network paths.
Then we watched how Adaptive Connect responded.
The experiment exposed two important flaws in our recovery logic. We implemented fixes, but the subsequent device test did not conclusively exercise either fix.
That distinction is central to this report.
This article explains what we tested, what failed, what we changed, and what still needs to be verified.
The original research, including the raw measurements, is available at Colitu Lab.
1. The Engineering Problem: Connectivity Is Not a Boolean
Many networking applications treat connectivity as a relatively simple state:
- Connected
- Disconnected
Real network behavior is considerably more complicated.
A useful connection is not merely one that successfully negotiated a transport.
It must continue carrying application traffic.
Consider a simplified connection lifecycle:
CONNECTING
|
v
HANDSHAKE SUCCESSFUL
|
v
TUNNEL ESTABLISHED
|
v
APPLICATION TRAFFIC WORKING
|
v
TRANSPORT STALLS
|
v
SOCKET STILL APPEARS CONNECTED
|
v
APPLICATION TRAFFIC FAILS
The dangerous transition happens between the last three states.
The transport may remain established even though it is no longer useful.
If the application relies on connection status alone, it may never recognize the failure promptly.
The restrictive-network problem
Our test environment exhibited a particularly interesting behavior.
Several TCP-based VPN transports could establish connections successfully, but traffic stopped progressing approximately 30–50 seconds later.
Hysteria2, which uses UDP, remained functional under the observed baseline conditions.
This created an unusual recovery problem.
The fallback system had several transport options available, but most of them could not sustain traffic on that network.
A successful handshake was therefore not sufficient evidence that a fallback would work.
Why adding more protocols is not enough
Suppose a client supports four transport options:
Hysteria2
VLESS Reality
Trojan
VLESS XHTTP
At first glance, this appears to provide considerable redundancy.
But imagine the following network conditions:
Hysteria2 WORKING
VLESS Reality CONNECTS -> STALLS
Trojan CONNECTS -> STALLS
VLESS XHTTP CONNECTS -> STALLS
The client has four choices, but only one is useful.
If the working transport fails temporarily, blindly cycling through the other three may prolong the outage.
Even worse, a cooldown mechanism might prevent the client from returning to the only transport that can actually carry traffic.
Protocol diversity provides opportunities for recovery. It does not guarantee recovery.
This distinction shaped our fault-lab design.
2. Why We Built a Fault-Injection Lab
Traditional connectivity tests answer questions such as:
- Can the client connect?
- Does the server respond?
- Can the tunnel carry a request?
- What is the latency?
Those are useful measurements.
They are also insufficient for evaluating a system that claims to recover from failures.
We needed answers to harder questions.
What happens when the active protocol disappears?
How quickly does the client detect a tunnel that has silently stopped forwarding data?
Does a backup connection actually survive until it is needed?
If a network restriction is removed, how quickly does the client return to a working transport?
Can a recovery mechanism accidentally prevent its own recovery?
These questions are difficult to answer using ordinary uptime monitoring.
A server can remain healthy while a particular client's route to that server is unusable.
A connection can remain established while application traffic is broken.
And a failure may last only long enough to affect a user without appearing in coarse monitoring data.
We therefore designed a controlled fault-injection suite.
Experimental environment
| Parameter | Configuration |
|---|---|
| Application | Colitu for Android |
| Baseline version | 2.8.0 |
| Test-build version | 2.8.0+ac2fix |
| Device | VIA E30, Android 11 |
| Network | Residential Wi-Fi in Russia |
| Connection mode | Automatic |
| Servers in scope | 29 |
| Failure scenarios | 11 in the baseline suite |
| Fault duration | 90–120 seconds |
| Traffic probe | HTTP request through VPN every 5 seconds |
| Measurement | Probe success and interruption duration |
| Test date | October 10, 2026 |
We selected a real device rather than relying solely on a simulated VPN client.
That matters because actual devices generate background traffic, change connection states, and interact with the networking stack in ways that simplified test environments may not reproduce.
As we discovered, ordinary background traffic was enough to expose one of our bugs.
3. How the Fault Lab Works
The fault injector operates through controlled firewall rules on Colitu infrastructure.
For each scenario, rules are applied to the relevant servers, targeting the test connection's public IP address.
Every rule is tagged so that it can be identified and removed.
Server-side expiry timers provide an additional safety mechanism: faults are not intended to persist indefinitely if the controlling process fails.
The test device continues sending real HTTP requests through its VPN tunnel at five-second intervals.
The experiment records:
- Whether each request succeeds.
- The longest observed interruption during the fault and its recovery window.
- The first successful request after fault injection begins.
- The time until successful traffic is observed after the fault is removed.
The following diagram illustrates the measurement approach.
ANDROID TEST DEVICE
|
|
Adaptive Connect
|
VPN Tunnel
|
v
COLITU SERVERS
|
Fault Injector
|
+-------------+-------------+
| | |
Drop UDP Reset TCP Drop Server
| | |
+-------------+-------------+
|
v
HTTP PROBES
Every 5 sec
|
v
RESULT LOG
|
+----------+----------+
| | |
Success Outage Recovery
Rate Duration Time
The fault injector and the traffic probes are conceptually separate.
The injector changes network conditions.
The probe system measures the resulting user-visible connectivity.
This separation is important: a successful control operation does not necessarily mean that the intended application path was disrupted.
The experiment must verify both.
Eleven failure scenarios
The baseline suite covered the following conditions:
| Scenario | Injected failure |
|---|---|
| 1 | Block Hysteria2 on all servers |
| 2 | Block VLESS Reality |
| 3 | Block Trojan |
| 4 | Inject TCP resets into VLESS Reality |
| 5 | Freeze TCP flows after 16 KB |
| 6 | Block UDP across all servers |
| 7 | Introduce 5% packet loss |
| 8 | Introduce 20% packet loss |
| 9 | Attempt a selected-server outage |
| 10 | Cut the server carrying the active tunnel |
| 11 | Block VLESS XHTTP |
The suite contains several distinct failure classes.
Protocol blocking tests whether a client can recover when a transport family becomes unavailable.
TCP resets model abrupt connection termination.
The 16 KB freeze models a more subtle situation: an initial exchange succeeds, then forwarding stops.
Packet loss examines behavior under degraded transport conditions.
Server removal tests whether the client can recover when the problem is not limited to a specific protocol.
Not every injected fault is expected to interrupt active traffic.
If the user is connected through Hysteria2 and Trojan is blocked, the existing connection may remain entirely unaffected.
That is not a failed experiment.
It is an important distinction between injecting a fault and causing an outage.
4. Baseline Results: Android 2.8.0
We began with the released Android application running Adaptive Connect 2.0.
The results immediately revealed areas requiring investigation.
Overall results
| Metric | Baseline result |
|---|---|
| Total HTTP probes | 466 |
| Successful probes | 397 |
| Failed probes | 69 |
| Probe success rate | 85.19% |
| Longest observed outage | 100.2 seconds |
| Unrecovered scenarios | 0 |
A success rate of 85.19% is not particularly impressive in isolation.
But this was not a normal-usage benchmark.
The network was deliberately disrupted, repeatedly, using several different failure mechanisms.
The important question was not simply how many requests succeeded.
It was why particular interruptions lasted as long as they did.
Outage breakdown
| Injected failure | Longest observed outage |
|---|---|
| Hysteria2 blocked globally | 95.2 s |
| VLESS Reality blocked | 25.0 s |
| Trojan blocked | 0 s |
| VLESS Reality TCP resets | 0 s |
| TCP frozen after 16 KB | 0 s |
| UDP blocked globally | 100.2 s |
| 5% packet loss | 0 s |
| 20% packet loss | 0 s |
| First selected-server scenario | 0 s |
| Active server removed | 45.1 s |
| VLESS XHTTP blocked | 0 s |
The zero-second entries require careful interpretation.
They indicate that the measurement did not observe a sustained interruption under the corresponding fault.
They do not establish that a protocol or application is immune to that class of failure.
With five-second probes, sufficiently short interruptions can also be missed.
Two failures stood out:
The unusually slow return after UDP became available again.
The prolonged interruption when the active server disappeared.
We investigated both.
5. Failure #1: A Cooldown That Delayed Recovery
One of the most revealing results came from the global UDP block.
The injected fault lasted 120 seconds.
The longest observed interruption was approximately 100 seconds.
More importantly, after the fault was removed, traffic did not resume immediately.
Recovery took approximately 31 additional seconds.
This was particularly interesting because Hysteria2 had previously been the transport that worked on the network.
Why was the application not returning to a known-good transport promptly?
The ten-minute penalty
When Hysteria2 stalled in the middle of an established session, Adaptive Connect placed the transport at the back of that server's candidate list for ten minutes.
The intention was reasonable.
If a transport fails, avoid repeatedly selecting it and entering an endless failure loop.
However, the observed network conditions made that policy problematic.
Hysteria2 had been the only transport capable of sustaining traffic in the original environment.
The alternative TCP transports could establish connections, but subsequently stalled.
The client therefore spent time exploring alternatives that could not reliably carry traffic.
Meanwhile, the transport that had demonstrated actual connectivity remained penalized.
When failure memory becomes harmful
This illustrates a broader networking problem.
A recovery system needs memory.
Without it, the client may repeatedly choose the same failing route.
But memory can become harmful when it does not distinguish between different kinds of failure.
Consider the difference between:
- A transport that has never successfully worked on the current network.
- A transport that previously worked and then failed during a temporary restriction.
These situations should not necessarily receive identical penalties.
The first may indicate a fundamentally unsuitable route.
The second may indicate a temporary condition that has already disappeared.
Treating both equally can delay recovery.
The additional missed-probe problem
The logs exposed a second timing issue.
Adaptive Connect enforced a 60-second minimum interval between automatic switches.
That interval was intended to avoid excessive switching.
However, the missed-probe counter was being reset before the minimum interval was evaluated.
The consequence was subtle.
A tunnel could already be demonstrably broken, yet the client would wait for the switching interval.
After the interval ended, the client still required three more missed checks before initiating another switch.
In other words, the system could spend time rediscovering a failure it already knew about.
The changes we implemented
We adjusted the recovery logic in two ways.
1. Shorter penalties for previously successful transports
A transport that has already worked on the current network now receives a 90-second mid-session penalty in the test build.
Other transports retain the ten-minute penalty.
The intent is to preserve caution around historically unsuccessful paths while allowing previously functional transports to return sooner.
2. Miss counters continue through the switching interval
Missed checks are no longer discarded merely because the minimum switching interval has not yet elapsed.
When that interval expires, the client can act on accumulated failure evidence rather than restarting the detection process.
Conceptual illustration
The following pseudocode illustrates the policy difference. It is explanatory, not a copy of Colitu's production implementation.
def failure_penalty(transport, network):
if transport.previously_worked_on(network):
return 90 # seconds
return 600 # seconds
def should_switch(health, elapsed_since_switch):
if health.consecutive_misses < 3:
return False
if elapsed_since_switch < 60:
return False
return True
The key idea is that health.consecutive_misses remains meaningful while switching is temporarily restricted.
It should not be reset simply because the application is observing a minimum transition interval.
The deeper lesson
A cooldown is not a substitute for understanding the failure.
Recovery policy must balance the risk of retrying too aggressively against the risk of excluding the only usable path.
The correct answer depends on what the transport has demonstrated on the current network.
6. Failure #2: Our Warm Spare Was Dead When We Needed It
Adaptive Connect 2.0 includes a warm spare mechanism.
The basic design maintains a primary path and a backup path within the running VPN connection.
When the primary path stops working, traffic can move to the backup without requiring a full VPN disconnect-and-reconnect sequence.
At least, that is the intended behavior.
A simplified architecture looks like this:
DEVICE
|
v
VPN BALANCER
|
+---------+---------+
| |
v v
MAIN PATH WARM SPARE
| |
v v
Server A Server B
| |
+---------+---------+
|
v
INTERNET
The backup provides useful redundancy only if it remains healthy.
And this is where our implementation encountered a problem.
The observed failure
During the baseline test, we deliberately removed the server carrying the active tunnel.
The client experienced a 45-second interruption.
When we inspected the logs, the relevant backup path was already dead.
It had not been replaced.
The system had a warm spare mechanism, but it did not have a usable spare at the moment of failure.
Why the spare was not replaced
Replacing a spare requires reloading part of the underlying connection core.
That operation can cause a brief interruption.
We therefore designed the application to prefer an idle moment before replacing a dead spare.
The original threshold required less than 10 KB of traffic over a ten-second interval.
This was intended to protect active user sessions.
However, the test device was not truly idle.
Background synchronization, application traffic, and connectivity checks generated approximately 15–30 KB over each ten-second observation window.
This was enough to keep the replacement operation deferred.
And deferred.
And deferred again.
The application was trying to avoid disrupting traffic.
As a consequence, it indefinitely postponed restoring redundancy.
A liveness problem hidden inside a safety mechanism
This is a classic engineering tradeoff.
We wanted to preserve two properties:
Safety: Avoid interrupting useful foreground traffic during backup replacement.
Liveness: Eventually restore a healthy backup after the current one fails.
Our original idle rule favored safety so heavily that liveness was not guaranteed.
Under persistent low-volume background traffic, there might never be an eligible replacement window.
The backup could remain unavailable indefinitely.
Our revised replacement policy
The test build introduces a more flexible rule.
Initially, the application still prefers an idle moment.
However, after waiting two minutes, light traffic no longer prevents replacement.
The revised threshold allows replacement below 32 KB per ten-second window.
This is not a guarantee that every replacement will be invisible.
It is an attempt to balance the risk of a brief maintenance disruption against the greater risk of remaining indefinitely without a working backup.
Why this matters beyond VPN software
The same failure pattern can appear in many fault-tolerant systems.
Imagine a database replica that is never rebuilt because the primary system is never fully idle.
Or a service worker that is never replaced because the application continuously receives a small amount of traffic.
Or a load balancer that indefinitely delays restoring a failed backend.
In each case, a mechanism intended to reduce disruption can prevent the system from returning to a resilient state.
Redundancy is not only about having a backup. It is about continuously maintaining the conditions that make the backup usable.
That is the engineering principle this failure exposed.
7. A Separate Windows and Linux Finding: Network Memory Must Have Boundaries
Adaptive Connect maintains information about transport behavior on different networks.
That memory helps the application avoid repeating known failures.
For example, a protocol that fails on a restrictive mobile network may still work perfectly well on home Wi-Fi.
Those environments should not share the same recovery assumptions.
A separate Windows investigation exposed a problem with how previous network state could influence a newly joined network.
Information such as stalled transports and the last successful transport could persist into a different network context.
We changed the relevant Windows and Linux source logic so that failure history from the previous network is not applied to the new one.
These changes are covered by unit tests.
However, they were not validated by the Android fault-injection experiment described here.
That distinction matters.
The existence of unit tests supports confidence in specific implementation behavior.
It does not establish end-to-end reliability across every operating system or network transition.
A complete evaluation still requires device-level testing of Wi-Fi changes, mobile handovers, and other transitions.
8. The Second Run Looked Dramatically Better
After implementing the changes, we ran the suite again using a test build.
The top-level results were striking.
| Metric | Released 2.8.0 | Test build |
|---|---|---|
| HTTP probes | 466 | 526 |
| Successful probes | 397 | 518 |
| Success rate | 85.19% | 98.48% |
| Longest outage | 100.2 s | 35.1 s |
| Unrecovered scenarios | 0 | 0 |
Some of the individual failure results also appeared substantially better.
| Failure | Released app | Test build |
|---|---|---|
| Hysteria2 globally blocked | 95.2 s | 0 s |
| VLESS Reality blocked | 25.0 s | 0 s |
| UDP globally blocked | 100.2 s | 0 s |
| Active server removed | 45.1 s | 35.1 s |
It would be tempting to publish this table with a headline such as:
Our fixes increased connectivity from 85% to 98%.
That would be an unsupported conclusion.
And the reason is important.
The two runs did not start from equivalent conditions
During the second run, the device had settled on a different server.
On that server, a TCP-based transport was functioning.
The initial restrictive-network behavior was therefore not equivalent to the first run.
When Hysteria2 was blocked, the device did not need Hysteria2.
When UDP was blocked, the active TCP connection continued carrying traffic.
These faults therefore did not recreate the condition that exposed the original cooldown problem.
The improved numbers are real observations.
But they are not proof that the cooldown fix caused the improvement.
The warm spare fix was not exercised either
During the second run's active-server outage, the logs once again showed that the spare was unavailable.
The client experienced an interruption of approximately 35 seconds.
The revised spare-replacement mechanism did not get the opportunity to demonstrate the intended recovery behavior.
As a result, this experiment cannot verify that the revised idle threshold solves the original problem on the device.
One scenario was skipped
The second run also did not execute every intended server-removal scenario.
One additional active-server cut was skipped because the test tool could not identify the tunnel on a listed server during its capture window.
That limitation is recorded in the raw result file.
The second run therefore has a different number of executed fault episodes from the first run.
It is not a perfectly matched before-and-after comparison.
What we can legitimately conclude
The second run provides evidence that the test build continued functioning under the faults it actually encountered.
It did not reveal a new failure in those observed conditions.
However:
- The revised Hysteria2 recovery policy was not conclusively validated.
- The revised warm spare replacement policy was not conclusively validated.
- The two probe-success percentages cannot be attributed solely to the code changes.
- The second run is not a controlled A/B comparison.
A better benchmark result is not the same thing as verified causality.
This is perhaps the most important methodological lesson from the experiment.
9. We Also Found Problems in the Test Infrastructure
Fault injection tests application behavior.
But they also test the assumptions made by the people designing the experiments.
Our first attempts exposed several problems in the testing infrastructure itself.
Problem A: Incomplete server coverage
An early version of the suite targeted eight servers.
The client eventually moved to a server outside the target list.
Subsequent injected faults did not affect the active tunnel.
The test tool continued working, but the experiment no longer exercised the intended application path.
We expanded the target scope to all 29 servers the test device could select.
We also added logging to identify which server was carrying the tunnel before each fault.
This provides a basic but essential validation step:
Did the experiment actually affect the intended target?
Problem B: Public-IP targeting and NAT
The firewall rules targeted the test device's public IP address.
This seemed like a convenient way to isolate traffic belonging to a particular experiment.
However, multiple devices on a residential network can share the same public IP through NAT.
During testing, other computers in the same home also experienced connectivity interruptions.
This did not imply that unrelated Colitu customers elsewhere were affected.
It did reveal that public-IP isolation was not equivalent to individual-device isolation.
The operational procedure was adjusted to announce fault runs in advance and move the control workstation to a server excluded from the experiment.
This limitation should remain explicit in any description of test safety.
Problem C: Sequential fault application
When a failure must affect many servers, sequentially applying firewall rules introduces timing differences.
One server may start dropping packets before another.
If an operation hangs, the entire sequence may stall.
We updated the control mechanism to apply server-side changes in parallel.
The tooling also supports continuing from a selected fault after interruption.
This improves consistency and operational robustness.
It does not eliminate every source of timing uncertainty, but it makes the injected scenarios more controlled.
10. What This Experiment Taught Us About Failover Design
Beyond the individual bugs, the experiment revealed several general principles.
Principle 1: A working socket is not evidence of a working network
Health checks should verify actual application-level progress.
A handshake is useful evidence.
An open socket is useful evidence.
But neither establishes that payload traffic can continue flowing.
This is especially important when failures are silent rather than explicit.
Principle 2: Recovery policy must remember context
A route that has never worked should not necessarily be treated like one that worked successfully until a temporary outage.
Failure memory must reflect observed behavior on the relevant network.
Otherwise, a policy intended to prevent repeated failures may accidentally prolong them.
Principle 3: Backup availability is a continuously maintained property
A spare path is not useful simply because it existed at some point.
It must remain healthy and be replaced when necessary.
Backup maintenance requires its own monitoring and eventual recovery guarantees.
Principle 4: Avoid treating idle time as a guaranteed event
Real devices are rarely completely quiet.
Background synchronization, health checks, notifications, and system services may generate continuous low-volume traffic.
Maintenance operations that require perfect idleness can suffer indefinite postponement.
Principle 5: The active network path matters more than the protocol list
Blocking an unused protocol may have no effect on the user.
Blocking the only viable transport may produce an unavoidable outage.
A meaningful fault-injection report must record the active transport and server before and during each scenario.
Without that context, the results can be misleading.
Principle 6: Recovery cannot create a working route where none exists
During the baseline experiment, the network already exhibited widespread TCP transport freezes.
When Hysteria2 was also unavailable, the client could temporarily have no usable route.
No fallback policy can guarantee continued connectivity if every available path is unusable.
The engineering objective is therefore not to promise impossible availability.
It is to identify usable alternatives quickly, avoid unnecessary recovery delays, and restore traffic as soon as a viable path returns.
Principle 7: A successful test result must be interpreted against its starting state
Different network conditions can produce dramatically different results with identical software.
Controlled comparisons require equivalent initial state, including the active server, active transport, relevant failure memory, and injected fault.
Otherwise, a change in the environment can be mistaken for a change in software quality.
11. Designing the Next Experiment
The immediate priority is not producing another large success-rate figure.
It is verifying that the fixes actually address the failures they were designed to solve.
We see several important next steps.
Experiment A: Verify recovery after temporary UDP loss
The test should deliberately begin with Hysteria2 carrying traffic.
The sequence would be:
1. Establish a working Hysteria2 tunnel.
2. Verify continuous application traffic.
3. Block UDP to the relevant servers.
4. Confirm that Hysteria2 stops working.
5. Allow the client to enter its recovery process.
6. Remove the UDP restriction.
7. Measure the time until traffic resumes.
8. Inspect transport-selection and cooldown logs.
9. Repeat with the baseline and test builds.
The critical measurement is recovery behavior after the original usable route becomes available again.
We should compare the released version and revised version under equivalent conditions.
Repeated trials would help distinguish a consistent improvement from ordinary network variability.
Experiment B: Verify warm spare replacement
A separate test should target the spare-liveness problem directly.
The setup must ensure that the client initially has a healthy main connection and a healthy spare.
Then the experiment should deliberately make the spare unusable while the primary continues carrying traffic.
After that, the client should be observed under controlled background-traffic levels.
The essential questions are:
- Does the client recognize that the spare has failed?
- Does it attempt replacement?
- Can background traffic postpone that replacement indefinitely?
- Does the new threshold allow a replacement under light traffic?
- Is the replacement itself disruptive?
- Once a new spare is ready, does it carry traffic when the main path fails?
The last step is crucial.
A successful replacement log is not enough.
The replacement must demonstrate that it can perform its intended function.
Experiment C: Improve measurement resolution
Five-second HTTP probes provide a useful starting point.
They do not capture every short-lived interruption.
Future experiments could use more frequent probes and correlate them with transport-level events.
Useful metrics include:
| Metric | What it tells us |
|---|---|
| Failure detection latency | How quickly the client identifies a broken path |
| Failover latency | Time between detecting a failure and selecting an alternative |
| Traffic restoration time | When application traffic becomes usable again |
| Backup readiness | Whether a healthy spare exists when needed |
| Recovery p50 / p95 / p99 | Distribution of recovery times across repeated trials |
| Probe success ratio | Fraction of successful application requests |
| False-positive switches | Healthy paths abandoned unnecessarily |
| Switching frequency | Whether recovery policy creates excessive oscillation |
A single maximum outage is useful.
A distribution across repeated, controlled trials is far more informative.
Experiment D: Validate the security properties during recovery
Connectivity is only one part of VPN reliability.
A future security-focused experiment should also measure what happens when the tunnel fails.
Important questions include:
- Does the operating system ever send protected traffic outside the VPN?
- Does the kill switch behave correctly during transitions?
- Are DNS requests handled safely during recovery?
- Can changing servers expose unexpected routing behavior?
- Do failed connections leave temporary routing or firewall inconsistencies?
These measurements are not provided by the present fault-injection dataset.
They require their own test design and validation.
A system should not trade away its intended privacy guarantees merely to restore connectivity more quickly.
Experiment E: Expand platform coverage
The current experiment was performed on one Android 11 device.
That is valuable, but narrow.
A broader evaluation should eventually include:
- Android devices with different networking stacks.
- Windows under Wi-Fi and Ethernet transitions.
- Linux under changing routes and interfaces.
- iOS under mobile and Wi-Fi handovers.
- Different network providers and filtering environments.
Platform-specific behavior matters.
A recovery mechanism that succeeds in one operating system's networking environment may behave differently in another.
12. Reproducibility Is Part of the Engineering Work
One reason we publish Colitu Lab research is to make the underlying measurements available for inspection.
For this experiment, the raw baseline and test-build results are published as JSON.
You can inspect the exact probe totals, fault definitions, durations, and recorded outages.
Raw data:
Research and methodology:
- Original fault-injection research
- Colitu Lab methodology
- Adaptive Connect 2.0 technical report
- Colitu VPN Lab GitHub repository
The GitHub repository contains broader VPN protocol experimentation resources. It should not be treated as proof that every script used in this particular Adaptive Connect regression experiment is already independently reproducible from that repository.
Complete reproduction requires the experiment-specific scripts, compatible infrastructure, configuration, and documented execution procedure.
Publishing result files is useful transparency.
Making the complete experiment repeatable by independent engineers is a further step.
Both matter.
13. What We Are Not Claiming
Technical reports become less useful when their conclusions extend beyond the evidence.
So it is worth stating explicitly what this experiment does not establish.
We are not claiming that Adaptive Connect delivers 98.48% availability across restrictive networks.
That number describes successful HTTP probes during one specific test-build run.
We are not claiming that the fixes reduced outages from 100 seconds to 35 seconds.
The two runs did not begin from equivalent network conditions.
We are not claiming that the warm spare replacement fix is verified on Android.
The second device run did not exercise the intended replacement behavior.
We are not claiming that eleven injected faults represent every real-world network restriction.
Real interference can involve more complex conditions, including latency manipulation, DNS behavior, route instability, and protocol-specific inspection.
We are not claiming that the Windows and Linux changes have passed equivalent live-network tests.
The changes described here were covered by unit tests, while this particular fault lab targeted Android.
Finally, the fixes described in this report were present in a test build at the time of the experiment.
They should not be confused with independently verified behavior across every released Colitu client.
These limitations do not make the experiment useless.
They define what its results can responsibly support.
14. Why We Believe Testing Failure Is More Valuable Than Demonstrating Success
It is relatively easy to demonstrate a successful VPN connection.
Open the application.
Select a server.
Connect.
Load a website.
That verifies an important basic function.
But it tells us very little about resilience.
The harder questions only appear when something goes wrong.
What if the server disappears after the connection has been established?
What if a transport works for 30 seconds and then silently freezes?
What if all UDP traffic is blocked?
What if a backup is unavailable when it is needed?
What if the network recovers, but the application's own cooldown policy delays reconnection?
These are the situations in which recovery architecture is tested.
And they are also the situations in which seemingly reasonable engineering assumptions can fail.
Our fault lab found a ten-minute transport penalty that was poorly suited to a particular recovery scenario.
It found a backup maintenance rule that could postpone replacement indefinitely.
It found limitations in the test infrastructure itself.
None of those findings is especially flattering when presented as a product benchmark.
But they are useful engineering results.
They identify specific behavior that can be changed, measured, and tested again.
More importantly, they help turn vague claims about reliability into concrete, falsifiable expectations.
Instead of saying:
The application automatically recovers from network failures.
We can ask:
When Hysteria2 is the only working transport and UDP becomes available again, how quickly does the client restore application traffic?
Instead of saying:
The application maintains a backup tunnel.
We can ask:
After the backup fails under continuous low-volume traffic, does the client eventually replace it, and can that replacement take over when the primary path disappears?
Those are questions that a serious fault-injection experiment can answer.
And if it cannot answer them yet, the report should say so.
15. Where Adaptive Connect Goes From Here
Adaptive Connect 2.0 is an attempt to make VPN connectivity more responsive to the conditions users actually encounter.
Its mechanisms include transport selection, network-specific failure memory, application-level health monitoring, server fallback, and warm spare recovery.
Each mechanism addresses a different part of the problem.
Together, they create a more complex system.
And complexity introduces interactions that are difficult to predict from isolated unit tests.
A transport penalty interacts with failure detection.
Failure detection interacts with switching intervals.
Backup maintenance interacts with ordinary device traffic.
Network memory interacts with interface transitions.
A useful test strategy must examine those interactions rather than evaluating each component only in isolation.
That is the direction we want Colitu Lab to pursue.
The next meaningful milestone is not a larger marketing number.
It is a controlled experiment demonstrating that the identified failure conditions are handled correctly by the revised implementation.
From there, the same methodology can be extended to more devices, more networks, and more realistic combinations of faults.
The long-term objective is straightforward:
When a usable network path exists, the client should find it, recognize when it stops working, and recover without unnecessary delay or loss of the protections users expect from a VPN.
That is an engineering objective, not a universal availability guarantee.
And it is something we can continue testing.
Final Thoughts
We began this experiment with a simple idea:
Break the network deliberately and measure what happens.
The baseline run produced 397 successful HTTP probes out of 466.
The revised test build produced 518 successful probes out of 526.
But the most useful results were not those percentages.
They were the two bugs found in the recovery logic, the implementation changes made in response, and the realization that the second experiment did not actually verify those changes under equivalent conditions.
That leaves us with unfinished work.
The Hysteria2 recovery adjustment needs targeted device validation.
The warm spare replacement policy needs a test that first kills the spare and then verifies replacement and takeover.
The experiment needs stronger control over its starting state.
And the measurement process can become more precise.
This is what meaningful reliability engineering often looks like.
Not a perfectly clean graph.
Not an uninterrupted sequence of successful demonstrations.
But a cycle of designing systems, breaking assumptions, examining evidence, correcting behavior, and testing again.
At Colitu, we want to keep publishing that process, including the parts that do not yet work as intended.
Because a VPN should not merely know how to connect.
It should know when its connection has stopped being useful.
And its recovery mechanisms should be tested just as aggressively as the networks they are designed to survive.
Built for networks that fight back.
Colitu — Connection Liberty Tunnel
Explore the research: Colitu Lab
Read the technical architecture: Adaptive Connect 2.0
Explore the project: colitu.com
Open-source networking experiments: github.com/colitu/vpn-lab
This article is based on Colitu Lab's October 10, 2026 fault-injection investigation. Results describe the measured experimental conditions and should not be interpreted as universal performance or security guarantees.
Top comments (1)
tr.ee/dev-to
Some comments have been hidden by the post's author - find out more