DEV Community

Gabriel Mahia
Gabriel Mahia

Posted on

Why East Africa's Public Domain Agricultural Data Is the Most Underused AI Asset in the Region

East Africa has a precise and detailed agricultural data record covering soil types, rainfall patterns, crop varieties, market conditions, and production methods across the region's major farming areas.

It was written between 1906 and 1925.

The USDA Foreign Agriculture Service published systematic crop surveys across Kenya, Uganda, and Tanganyika during this period. The Royal Botanic Gardens at Kew published botanical surveys throughout East Africa. Colonial agricultural departments published monthly bulletins covering rainfall, market prices, planting calendars, and extension notes.

All of it is public domain. None of it is structured for AI access.

What's in the record

The USDA surveys from 1910-1920 document varietal performance by altitude and rainfall zone, soil type classifications by district, market price series for maize and other staples, planting and harvest calendars, and pest and disease records.

These records predate synthetic fertilizers and irrigation as primary inputs. They document how crops performed under natural rainfall conditions across the same climate zones that East African smallholder farmers still farm today. For AI-assisted crop advisory without irrigation infrastructure — which describes most East African smallholder farming — this is directly relevant baseline data.

The structuring problem

The data is not in a format an AI agent can query. It's in PDFs, some scanned from physical documents. It's in library archives. Structuring it for AI access is the missing work. Not generating it — it already exists.

The east-africa-agricultural-pd dataset at huggingface.co/datasets/gmahia is the beginning of that structuring work. It feeds kilimo-mcp, wapimaji-mcp, and bima-mcp.

The constraint is not data availability. It's the engineering work of converting what exists into what AI agents can use.

Top comments (0)