A rotating residential pool with 1,000 IPs sounds like abundance until you measure it. On any given hour, 4-7% of the pool is broken: unreachable, geo-drifted, ISP mislabeled, or already on someone else's abuse list. Nobody tells you which ones. Your account gets flagged, you check the proxy provider dashboard, and it shows "all systems operational."
The only fix is your own monitoring stack. This article covers the one that has held up across three production pools: what to check, how often, where to store the data, which Grafana panels actually catch incidents before users do, and how the health signal loops back into the BitBrowser profile fleet so bad IPs stop getting reassigned.
What "proxy health" actually means
Proxy health is not a single metric. It is a rolling composite of five signals, each of which fails independently:
- Connectivity. Does the proxy answer a TCP handshake in under 3 seconds?
- Egress IP correctness. When you route a request through it, does the destination see the IP the provider promised?
- Geolocation accuracy. Does the IP resolve to the country and region the provider claims?
-
ISP classification. Does the IP register as
residential(ormobile) rather thanhostingordatacenter? - Reputation status. Is the IP on Spamhaus DROP, on Scamalytics' fraud list, or blocked by Cloudflare's threat score?
A proxy can pass 1-3 and fail 4-5, and the failure is invisible until an account signup gets a CAPTCHA loop or a payment gateway declines the transaction. All five need continuous checks.
The metrics catalog
The dashboard tracks eleven metrics per IP. Nine come from active probes, two from the fleet feedback loop:
-
proxy_up(0 or 1) -
proxy_latency_ms(histogram) -
proxy_egress_ip_match(0 or 1) -
proxy_geo_country,proxy_geo_region(labels) -
proxy_isp_type(residential, mobile, hosting, unknown) -
proxy_reputation_score(0-100, lower is worse) -
proxy_blocklist_hits(count of lists flagging the IP) proxy_last_check_timestamp-
proxy_bitbrowser_profile_flags_24h(from the fleet feedback loop) -
proxy_bitbrowser_session_failures_24h(from the fleet feedback loop)
Eleven metrics, one label set per IP, roughly 11,000 time series for a 1,000-IP pool. Prometheus handles that on a t3.medium without breaking a sweat.
The architecture
Proxy pool (Proxy-Seller, IPRoyal, Bright Data, ...)
│
▼
┌───────────────────────────────┐
│ Check worker (Node.js) │
│ ─ Round-robin scheduler │
│ ─ HTTP probe (via proxy) │
│ ─ Geo lookup (ipinfo.io) │
│ ─ Reputation lookup │
│ ─ Prometheus /metrics │
└───────────────┬───────────────┘
│ scrape /metrics every 15s
▼
┌───────────────────────────────┐
│ Prometheus │
│ ─ 30d retention │
│ ─ Alertmanager rules │
└───────────────┬───────────────┘
│
▼
┌───────────────┴───────────────┐
│ Grafana │
│ ─ Pool overview │
│ ─ IP drill-down │
│ ─ Provider comparison │
└────────────────────────────────┘
One check worker per 500 IPs. Two workers cover a 1,000-IP pool with room to spare. Prometheus scrapes both workers on the same interval so metrics stay aligned.
The check worker
Node.js with axios, prom-client, and a small round-robin queue. The core loop for one proxy:
const axios = require('axios');
const { HttpsProxyAgent } = require('https-proxy-agent');
const client = require('prom-client');
const proxyUp = new client.Gauge({
name: 'proxy_up',
help: '1 if the proxy answered, 0 otherwise',
labelNames: ['proxy_id', 'provider', 'geo_expected']
});
const proxyLatency = new client.Histogram({
name: 'proxy_latency_ms',
help: 'Round-trip latency to the probe endpoint',
labelNames: ['proxy_id', 'provider'],
buckets: [50, 100, 250, 500, 1000, 2500, 5000]
});
async function checkOne(proxy) {
const agent = new HttpsProxyAgent(proxy.url);
const start = Date.now();
const labels = { proxy_id: proxy.id, provider: proxy.provider, geo_expected: proxy.geo };
try {
const res = await axios.get('https://ipinfo.io/json', {
httpsAgent: agent,
timeout: 8000
});
const elapsed = Date.now() - start;
proxyUp.set(labels, 1);
proxyLatency.observe({ proxy_id: proxy.id, provider: proxy.provider }, elapsed);
return { ok: true, data: res.data, elapsed };
} catch (err) {
proxyUp.set(labels, 0);
return { ok: false, error: err.code || err.message };
}
}
The probe endpoint is ipinfo.io/json because it returns IP, country, region, city, and ASN in a single call. Free tier gives 50,000 requests per month, enough for a 1,000-IP pool checked every 5 minutes.
Reputation lookups are more expensive. Batch them: pull every distinct IP from the pool once per hour, hit IPQualityScore or Scamalytics in bulk mode, cache results for 60 minutes in Redis. This keeps the reputation API bill under $30/month even at 1,000 IPs.
Storage choice
Prometheus over InfluxDB, TimescaleDB, or Victoria Metrics for a pool this size. Reasons:
- The metric cardinality (11 metrics × 1,000 IPs = 11k series) sits comfortably in the middle of Prometheus's sweet spot
- Grafana's Prometheus data source is the most polished
- The Prometheus query language handles the "IPs currently dead" and "IPs that failed 3 checks in a row" queries in one line each
- Alertmanager ships with Prometheus and needs no extra binary
Retention: 30 days. Anything older than a month is not useful for triage and takes up disk. If you need long-term trends, export daily aggregates to a Postgres table and query that instead.
The Grafana dashboard
Four panels. That is it. Adding more panels reduces incident response speed, not increases it.
Panel 1: Pool health heatmap. X-axis is time (last 6 hours). Y-axis is proxy_id. Cell color is proxy_up. A sudden vertical stripe of red means a provider outage. A slowly growing pattern of red dots means IPs are being burned individually.
Panel 2: Failed-probe rate per provider. Line graph, one line per provider (Proxy-Seller, IPRoyal, NetNut, DataImpulse). Y-axis is (1 - avg(proxy_up)) * 100. Any line crossing 8% is a signal to open a support ticket with that provider.
Panel 3: Geo drift table. A table of IPs where proxy_geo_country != geo_expected. If a US IP starts resolving to Panama, remove it from the pool immediately.
Panel 4: Correlation with BitBrowser flags. Scatter plot: X-axis is proxy_reputation_score, Y-axis is proxy_bitbrowser_profile_flags_24h. The cluster of high-flag / low-reputation IPs in the bottom-right is your daily kill list.
Alertmanager rules
Three alerts, no more. Every additional alert dilutes the ones that matter:
groups:
- name: proxy-health
rules:
- alert: ProxyPoolDegraded
expr: (1 - avg by (provider) (proxy_up)) > 0.08
for: 10m
labels:
severity: warning
annotations:
summary: "{{ $labels.provider }} pool is >8% down for 10 minutes"
- alert: ProxyGeoDrift
expr: count(proxy_geo_country != on(proxy_id) group_left proxy_geo_expected) > 20
for: 5m
labels:
severity: warning
- alert: ProxyReputationCollapse
expr: avg by (provider) (proxy_reputation_score) < 40
for: 15m
labels:
severity: critical
Route them to a dedicated Slack channel that only fires on proxy issues. Muted-by-default. Anyone whose account gets flagged checks that channel first.
Feeding the signal back into BitBrowser
The dashboard is not the endpoint. It is the input to a decision loop. A small service (also Node.js, also under 300 lines) subscribes to Prometheus queries every 5 minutes and, for any IP flagged as unhealthy:
- Calls the BitBrowser LocalAPI to find every profile bound to that IP
- Marks each profile as
proxy_stalein a local database - Prevents the dispatcher from opening those profiles until a new IP is assigned
- Requests a fresh IP from the provider's rotation API
- Updates the profile via
/browser/updatewith the new proxy config
This closes the loop. A dying IP takes profiles offline in under 5 minutes instead of after a user complaint. Protocol-level nuances (why HTTP vs SOCKS5 vs Shadowsocks matter here, especially for mobile carrier detection) are covered in this proxy protocol comparison.
Where cloud phones fit in the same pipeline
Mobile pools have a different failure mode: the phone itself can be fine while the carrier's IP block goes on Snapchat's ban list. Extending the dashboard to cover BitCloudPhone Android and BitCloudPhone iOS devices means adding two more probe types:
- App-store reachability check. Every 15 minutes, from each cloud phone, hit the App Store or Play Store discovery endpoint. A 403 or 429 means the device's carrier IP is soft-blocked at the app layer, invisible from web-only probes.
- In-app CAPTCHA rate. Track how often accounts running on a given device hit CAPTCHAs. Above a threshold, cycle the device.
Both probes return the same Prometheus metric names with a platform=ios or platform=android label. One dashboard covers three worlds.
Scaling notes
The stack above sizes to roughly 5,000 IPs on a single Prometheus instance and a two-worker check fleet. Beyond that:
| Layer | Bottleneck | Fix |
|---|---|---|
| Check worker | Node event loop under 2,000 concurrent HTTP calls | Split into worker pool via piscina
|
| Prometheus | Series cardinality above 100k | Shard by provider onto separate Prometheus instances |
| Reputation API | Cost per lookup | Extend cache TTL, batch harder, drop reputation for internal-only pools |
| Grafana | Panel query time above 3 seconds | Add recording rules for expensive expressions |
At 10,000 IPs the architecture pattern is the same, the numbers on each box change.
FAQ
Which reputation service is worth paying for?
IPQualityScore for general fraud scoring, Scamalytics for stricter e-commerce use cases, and MaxMind minFraud when payment risk is the specific concern. All three have similar accuracy; the choice comes down to which one your target platforms use themselves.
Can I skip reputation checks entirely?
For internal automation that never touches a payment gateway or an identity-verification vendor, yes. For anything that touches Stripe, Facebook Ads, or account creation on a Tier-1 platform, no. The reputation check catches issues 24-48 hours before the destination platform's own risk engine does.
What if my proxy provider does not expose sticky sessions?
The dashboard still works, but the metric labeling shifts. Instead of proxy_id, label by session_token and expect a much higher series churn. Prometheus handles it up to about 20,000 unique tokens per hour.
How do I probe from the same geo as the proxy claims to be in?
For most checks the probe origin does not matter — ipinfo.io returns the same data whichever direction the query comes from. For app-store and CAPTCHA checks, the probe must originate from the proxy itself; that is what the cloud phone probes cover.
Does this replace the proxy provider's own dashboard?
Yes, and you should stop trusting the provider's dashboard the day you have your own. Provider dashboards report internal health. What matters is what the destination platforms see, and only your own probes measure that.
Where to take it from here
Start with 50 IPs from one provider and the four Grafana panels. Prove the check worker stays under 10% CPU. Prove the alert rules fire on real outages, not on transient blips. Then add the second provider, then the third, then the cloud phone probes.
Every monitoring stack I have shipped in production started as one worker checking one pool with two panels. The stack that tried to cover everything from day one is still in a Git branch nobody merged.
Affiliate disclosure: This post contains affiliate links. If you sign up for a paid BitBrowser or BitCloudPhone plan through the links above, I may earn a small commission at no additional cost to you. All numbers above come from monitoring stacks I run against my own paid accounts.
Top comments (0)