Your nginx access log format can carry the client's country, city, and ASN. The lookup happens at write time, against a local database file, with no API call in the request path.
Most writing on this treats it as a module install. That is the wrong frame. The two modules worth using read the same file format, so the module is plumbing. The database you load is what decides which fields exist.
TL;DR
- Two nginx modules, one file format. Both read MMDB through
libmaxminddb. - The database decides which fields you get. GeoLite2 covers country, city, and the ASN number and organization, and it is free.
- ASN type, company, hosting provider, and threat flags need a richer database.
- Quote every variable in a JSON log format. An unquoted empty value writes
"threat":,and hard-fails the parser. I tested this. - Fix the client IP before you log it. One module trusts a client-controlled header by default.
- Swap database files with
mv, nevercp. The module has the file memory mapped. Pick the database for the fields you actually need, point whichever module you already have at it, then quote everything in the log format. The failure modes are all in the last two steps, not the install.
Which database gives you which field
This is the table I wanted when I started and could not find anywhere.
| Field | GeoLite2 City | GeoLite2 ASN | IPGeolocation Geo | IPGeolocation ASN | IPGeolocation Security |
|---|---|---|---|---|---|
| Country code | yes | no | yes | no | no |
| City, latitude, longitude | yes | no | yes | no | no |
| Currency, timezone, calling code | no | no | yes | no | no |
| ASN number | no | yes | no | yes | no |
| ASN organization | no | yes | no | yes | no |
| ASN type (ISP, HOSTING) | no | no | no | yes | no |
| ASN RIR, allocation date | no | no | no | yes | no |
| Company name, domain, type | no | no | no | no | separate Company DB |
| Hosting provider | no | no | no | no | separate Hosting DB |
| Threat score, VPN, proxy, Tor flags | no | no | no | no | yes |
Two things to know before you build a log format around this.
Everything arrives as a string. A threat score of 80 is the text 80, coordinates are text, and booleans come through as 1 or 0. Nginx variables have no other type, so whatever consumes the log does the casting.
ASN number formatting is not consistent across sources. GeoLite2 stores autonomous_system_number as an integer. Ours is documented as a bare integer in the database field reference, while the module docs show the variable rendering as AS1257. That difference decides whether your parser needs to strip a prefix, so check your own file rather than trusting either of us.
Two modules, one file format
Both candidates read MMDB files through libmaxminddb. That is the whole reason the database is the real decision.
| ngx_http_geoip2_module | ngx_http_ipgeolocation_module | |
|---|---|---|
| Install | Distro package on most systems | Build nginx from source |
| Field access | You write the path mapping yourself | Ready-made $ip_* variables |
| Multiple databases | One block per file | Repeat the directive, first match wins |
| Reload on update | auto_reload |
Reload nginx |
| Stream module | Yes, with --with-stream
|
HTTP only |
On Ubuntu 24.04 the first one is one command:
# Pulls libmaxminddb as a dependency and drops a load_module line
# into /etc/nginx/modules-enabled/ for you.
sudo apt-get install -y libnginx-mod-http-geoip2
That installed cleanly for me at version 1:3.4-5build2 alongside nginx 1.24.0.
Pick by what you already have. If the packaged module is there, keep it and write the mappings. If you are compiling nginx anyway and want ASN type or threat flags without hand-writing a path for every field, the ready-made variables save real config. Rebuilding a working nginx purely to avoid writing ten mapping lines is not a trade I would make.
There is a third module, the original ngx_http_geoip_module in the nginx tree. It reads MaxMind's retired legacy format and is not built by default. Do not start anything new on it.
Recipe A: map the fields yourself
The packaged module takes a file and a set of paths into it. Each path is a sequence of keys, space separated, and the module's directive reference covers the rest.
load_module modules/ngx_http_geoip2_module.so; # skip if your package enabled it
http {
geoip2 /etc/nginx/ipgeo/GeoLite2-City.mmdb {
auto_reload 5m; # re-reads the file without a reload
$geo_country default=ZZ country iso_code;
$geo_city city names en;
}
geoip2 /etc/nginx/ipgeo/db-ip-asn.mmdb {
auto_reload 5m;
$geo_asn asn as_number;
$geo_asn_org asn organization;
$geo_asn_type asn type; # not present in GeoLite2-ASN
}
}
The paths differ per file because the schemas differ, and that is the only thing you change when you swap databases. GeoLite2 nests country data under country and the ASN fields at the top level as autonomous_system_number and autonomous_system_organization. Ours nests everything under asn.
I checked that the packaged module handles an arbitrary nested schema rather than only MaxMind's own, because Recipe A falls apart otherwise. Building a small MMDB with the asn nesting above and pointing the stock Ubuntu module at it returned asn=15169 org=Google LLC type=BUSINESS on a hit. A miss and an RFC 1918 address both came back as empty strings for every variable, with no error and no refusal to start.
Tip:
default=only fires when the lookup misses, sodefault=ZZgives you a sentinel country you can filter on later instead of a silent empty string. Useful on the country field, noise everywhere else.
Recipe B: use ready-made variables
The other module skips the mapping. You load files, it exposes variables.
http {
# Declaration order is priority order: the first database
# containing a requested field answers for it.
ipgeolocation_db /etc/nginx/ipgeo/db-ip-geolocation.mmdb;
ipgeolocation_db /etc/nginx/ipgeo/db-ip-asn.mmdb;
ipgeolocation_db /etc/nginx/ipgeo/db-ip-security.mmdb;
# Default is `on`, which reads the leftmost X-Forwarded-For value.
# See the client IP section below before you leave it that way.
ipgeolocation_trust_forwarded_header off;
}
That gives you $ip_country_code, $ip_city_name, $ip_asn, $ip_asn_organization, $ip_asn_type, $ip_threat_score and about seventy more, with no paths to maintain. Booleans normalise to 1 and 0, and array fields like VPN provider names join into one comma-separated string. Build steps and the full variable list are in the module's build and variable docs.
The cost is the build. It compiles into nginx against libmaxminddb, and a dynamic build has to match your running binary exactly or nginx rejects it as not binary compatible. Reuse the flags from nginx -V and add --with-http_realip_module, which is not built by default and which you will want in a moment.
One genuinely nice property: if a database fails to open, nginx refuses to start. A typo in a path is caught at deploy rather than becoming empty columns you notice three weeks later.
A JSON access log format that survives empty values
An nginx log_format is a string template with no idea what JSON is. That is where these setups break, and it has nothing to do with geolocation.
log_format geo_json escape=json '{'
'"time":"$time_iso8601","ip":"$remote_addr","status":"$status",'
'"country":"$geo_country","city":"$geo_city",'
'"asn":"$geo_asn","asn_org":"$geo_asn_org","asn_type":"$geo_asn_type"'
'}';
access_log /var/log/nginx/access.geo.log geo_json;
Every value is quoted, including the ones that look numeric. That is deliberate. Private addresses, unlisted ranges, and any miss resolve to an empty string, and an unquoted empty value produces this:
{"ip":"127.0.0.1","threat":,"ok":"US"}
That is not valid JSON. I ran it through a parser to be sure, and it fails outright at the empty value. Quoted, the same miss writes "threat":"", which parses and lets you decide what an empty score means downstream.
escape=json handles the other half. City names and organisation names contain quotes, backslashes and non-ASCII characters, and without it one Turkish city name corrupts the line for every parser reading that file.
One subtlety worth knowing, because the nginx log module docs blur it. Their note that a hyphen is logged when the variable value is not found applies to variables that are genuinely unset, like $upstream_addr on a request that never reached an upstream. A geo module on a miss gives you a variable that exists and is empty, which is a different case and logs as empty. I tested both under each escape mode. The practical upshot: under escape=default an unset variable becomes the string -, which sails through a JSON parser and then poisons anything casting that field to a number. Under escape=json you get "" instead.
Get the client IP right before you log it
If nginx sits behind a load balancer, everything above geolocates your load balancer.
The fix differs by module, and one of the defaults is dangerous.
| Module | Default source | How to correct it |
|---|---|---|
ngx_http_geoip2_module |
$remote_addr |
Set geoip2_proxy for trusted ranges, or source=$variable
|
ngx_http_ipgeolocation_module |
Leftmost X-Forwarded-For
|
Set ipgeolocation_trust_forwarded_header off and use realip |
The second default is the one to watch. The leftmost X-Forwarded-For value is written by the client. Plenty of proxies append to that header instead of replacing it, so a visitor can send their own and choose which country they appear from. If your log is decorative, that costs you nothing. If anything reads it to make a decision, it is a hole.
http {
ipgeolocation_trust_forwarded_header off; # ignore the client-supplied header
set_real_ip_from 10.0.0.0/8; # your load balancer CIDR, nothing wider
real_ip_header X-Forwarded-For;
real_ip_recursive on; # walk back past trusted hops
}
The realip module rewrites $remote_addr to the real client before the geo lookup runs, using only hops you named. Both modules then geolocate the right address. Get the trusted range wrong and you are back to trusting the client, so keep it tight.
Derive fields at write time
Once the variables exist you can compute things before the line is written, which beats doing it in every query later.
# ASN type casing varies between sources, so match case-insensitively.
map $geo_asn_type $is_hosting {
default 0;
"~*hosting" 1;
"~*business" 0;
}
# Strip an AS prefix if your database renders one, so the field is
# always a bare number for whatever parses the log.
map $geo_asn $asn_num {
default $geo_asn;
"~^AS(?<n>\d+)$" $n;
}
# Datacenter traffic into its own file. A request is skipped when the
# condition is 0 or empty, so misses stay in the main log.
access_log /var/log/nginx/access.hosting.log geo_json if=$is_hosting;
Splitting hosting traffic out is useful because ASN type tells you whether an address belongs to a hosting network. It does not tell you whether the request is human or automated, so treat it as context rather than identity.
Keeping databases fresh without dropping requests
Both modules memory map the file. Overwriting it in place is how you get workers reading half a file.
#!/bin/sh
set -e
DIR=/etc/nginx/ipgeo
curl -fsSL "$DOWNLOAD_URL" -o "$DIR/db-ip-asn.mmdb.tmp"
# Same filesystem, so mv is atomic. cp and wget -O are not.
mv "$DIR/db-ip-asn.mmdb.tmp" "$DIR/db-ip-asn.mmdb"
# Never reload onto a truncated download.
nginx -t && nginx -s reload
With auto_reload set, the packaged module picks the new file up on its own and you can drop the reload line. Without it, reload. Either way the atomic mv is not optional.
If you are on GeoLite2, the licence terms require you to keep the data current, and geoipupdate is the supported path. Update cadence varies by database: IPGeolocation publishes refreshed databases daily, while MaxMind currently publishes GeoLite ASN every weekday and GeoLite City and Country on Tuesdays and Fridays.
Write time, ingest time, or query time
Doing this in nginx is one of three places, and it is not always the right one.
| Write time (nginx) | Ingest time (pipeline) | Query time (dashboard) | |
|---|---|---|---|
| Request latency | none, local lookup | none | none |
| Database lives on | every web node | one pipeline node | one place |
| Re-enrich old logs | not possible | reprocess | free |
| Reflects | the moment of the request | the moment of ingest | today's data |
The trade is historical accuracy against flexibility. A write-time country is what that address resolved to when the request happened, which is what you want for an incident review and wrong when an address changes hands and you need to re-attribute.
My rule: enrich at write time when nginx already needs the data to make a decision, because the lookup is happening anyway and the log line is free. Otherwise enrich at ingest, where one node holds the database and you can reprocess. If you are running the Elastic Stack there is a companion piece on enriching at ingest instead, and if you want these fields on a map there is one on shipping them to Loki and plotting them in Grafana.
What to do with the ASN column
A country column gets looked at. An ASN column mostly does not, which is a waste, because it is the field that tells you who actually owns the traffic.
Start by ranking:
jq -r '.asn_org // empty' /var/log/nginx/access.geo.log \
| sort | uniq -c | sort -rn | head -20
Your top twenty organisations usually cover most of your traffic. That is a small enough set to enrich properly, once, instead of per request:
import os
import sys
import requests
API_KEY = os.environ.get("IPGEO_API_KEY")
ASN_URL = "https://api.ipgeolocation.io/v3/asn"
def describe(asn):
"""Fetch routing context for one ASN. Returns None on any failure."""
if not API_KEY:
raise RuntimeError("IPGEO_API_KEY is not set")
try:
resp = requests.get(
ASN_URL,
params={
"apiKey": API_KEY,
"asn": asn,
# Relationships cost nothing extra; the charge is per ASN.
"include": "upstreams,peers",
},
timeout=(2.0, 10.0),
)
resp.raise_for_status()
except requests.RequestException as exc:
print(f"lookup failed for {asn}: {exc}", file=sys.stderr)
return None
data = (resp.json() or {}).get("asn") or {}
return {
"number": data.get("as_number"),
"org": data.get("organization"),
"type": data.get("type"),
"rir": data.get("rir"),
"ipv4_routes": data.get("num_of_ipv4_routes"),
"upstreams": [u.get("as_number") for u in data.get("upstreams") or []],
}
Twenty organisations is twenty lookups, once, cached in whatever you already run. The response carries more than the database does, including the official ASN handle, allocation status, announced route counts, and the upstream and peer relationships:
{
"asn": {
"as_number": "AS24940",
"organization": "Hetzner Online GmbH",
"country": "DE",
"type": "HOSTING",
"domain": "hetzner.com",
"date_allocated": "2002-06-03",
"asn_name": "HETZNER-AS",
"allocation_status": "ASSIGNED",
"num_of_ipv4_routes": "91",
"num_of_ipv6_routes": "5",
"rir": "RIPE"
}
}
Note the route counts and as_number come back as strings here too. Consistency, at least.
A few extra notes
IPv6. Both modules look up whichever address the client connects with, so the only thing that usually goes wrong is nginx not listening on v6 at all. Check for listen [::]:80;.
Private addresses. Internal health checks and anything behind a NAT will log empty geo fields forever. That is correct behaviour, not a misconfiguration, and it is the main reason the quoting rule above matters.
Retention. An IP address can be personal data, and in access logs where it can reasonably be linked to a visitor it should generally be treated that way. Adding city-level geolocation is additional processing. Enriching at write time means that data lands in a file that may have a longer retention period than your application database, so check the retention policy before turning this on fleet-wide.
Stream traffic. The packaged module works in stream blocks if you compiled with --with-stream, so TCP proxying can carry the same data. The other module is HTTP only.
Decide what an empty value means to you before you write a single if against these variables, because nginx treats both 0 and empty as false and that turns a missing database into a silently permissive rule. And pin your module build to your nginx version in CI. The one that bites is a routine package upgrade moving nginx out from under a dynamic module that was compiled against the old one.
Top comments (1)
One safeguard I'd add: never turn HOSTING into a deny rule by itself. Cloud ranges are often mislabeled as proxies. I require agreement from two dedicated proxy sources for hosting IPs, while a known anonymity-provider ASN can stand on its own.