About¶
Groundfact is a governed semantic layer over official statistics: dozens of public sources consolidated into one model you can question in plain English, where every number on screen carries its source, dataset and snapshot date.
Nepal is the first instance. Nothing in the architecture is specific to it.
It is built by Binit Ojha / TWAKKA LTD, in the open, as a working demonstration rather than a prototype — the site you are reading is the system, serving live from the same build this documentation was rendered from.
Why it exists¶
Ask a straightforward question about a country's development — how has child mortality moved,
what does school enrolment look like now — and the answer is genuinely hard to assemble. The
figures exist, publicly and freely licensed, but they arrive from a dozen institutions in a
dozen shapes, on different schedules, under different licences, using different identifiers
for the same country. Nepal is NPL to the World Bank, 524 to the UN, 149 to FAOSTAT and
NP to the DHS Program.
Worse, they disagree. The World Bank publishes modelled estimates; the DHS publishes survey results; a census publishes an enumeration. All three are legitimate, they differ for real methodological reasons, and almost nothing in the public tooling shows you where or why.
The usual response to this is a dashboard, which fixes the presentation and none of the governance. Groundfact is an attempt at the other order: fix the model first, then let the interface be thin.
What is actually in it¶
Read from the current build, not a brochure:
| Sources | 9, each with its licence recorded and CI-enforced |
| Observations | 646,483 |
| Governed observations | 511,054 resolved to a curated indicator |
| Curated indicators | 52 |
| Geographies | 260 |
| Cross-source comparisons | 43,175 place–period pairs where two sources both report |
Those last two rows are the interesting ones. 43,175 pairs where independent official sources report the same measure for the same place and year is not a byproduct — it is the thing worth building.
What makes it different¶
Disagreement is a feature, not an error. Where two sources report the same indicator for the same place and period and differ, Groundfact shows both figures, the gap, and a curated explanation of why they differ — rather than silently picking a winner or averaging them. Nepal's 2008 unemployment rate is 1.38% by ILOSTAT's labour force survey and 10.6% by the World Bank's modelled series. Neither is wrong. The survey ran in specific years under a definition that changed in 2017; the modelled series is smoothed and annual. Showing that, with the reason attached, is more useful than resolving it out of sight.
The language model never writes SQL and never produces a number. It resolves a question into a small validated specification — which indicators, which places, which years — using only identifiers that already exist. Deterministic code turns that into a parameterised query; the server binds the real values into the chart. A model can propose that a line for Bangladesh exists, but not what is in it. How a chart is made traces this with a real run, including the actual SQL and its bound parameters.
Provenance is attached to the number, not the page. Each row carries the source, dataset and snapshot date it came from, and those travel with it into the citation chip beside the figure. It is not a footnote assembled afterwards.
Reproducible as of a date. Every ingestion writes an immutable, dated snapshot that is never modified. A number cited on a given date is reproducible from the file that produced it, not from whatever the upstream API returns today.
The documentation cannot drift from the data. One build step renders the indicator definitions into this documentation, into the grounding document the model is given, and into the warehouse's own indicator dimension. A CI check re-renders all three and fails if any differs from what is committed. The claim "the docs match the system" is enforced rather than maintained.
What it is not¶
It is not a dashboard product, it is not a data marketplace, and it is not trying to be a general-purpose analytics tool. It answers questions about curated indicators from official statistical sources, cited, and says so plainly when it cannot.
It also does not have an agent framework, a vector database it does not need, or a text-to-SQL layer. Those were considered and rejected on the merits, and the reasoning is written down rather than implied.
For a short post¶
Lift-ready copy, factual and checkable.
One line
Groundfact — ask a country's official statistics a question in plain English, and get an answer where every number carries its source, dataset and snapshot date.
Short (LinkedIn-length)
I've been building Groundfact: a governed semantic layer over official statistics.
The problem it solves is boring and real. The same country's numbers arrive from a dozen institutions in a dozen shapes, using different codes for the same place, and they disagree for legitimate methodological reasons that nothing surfaces. Nepal's 2008 unemployment rate is 1.38% by ILOSTAT's survey and 10.6% by the World Bank's modelled series. Both are right.
So Groundfact consolidates 9 sources into one conformed model — 646,483 observations, 52 curated indicators, 260 geographies — and surfaces 43,175 place–year pairs where two official sources both report, showing the gap and the reason for it rather than picking a winner.
The design constraint I'm most pleased with: the language model never writes SQL and never produces a number. It resolves a question into a validated specification using only identifiers that already exist; deterministic code builds the query; the server binds the values into the chart. A model can propose that a line exists, but not what's in it. That makes fabrication structurally impossible rather than merely discouraged — and it makes prompt injection uninteresting, because the widest blast radius is picking a different valid indicator.
No agent framework, no vector store, no text-to-SQL. About 1,500 lines of Python for the whole AI layer. Nepal is the first instance; the architecture is country-agnostic.
Built in the open: groundfact.com
If you want the technical angle instead
Everything interesting about putting an LLM in front of a database is what you don't let it do.
In Groundfact the model gets a ~4,800-token grounding document listing every indicator and geography that exists, and returns one JSON object: which indicators, which places, which years. 131 output tokens. That gets validated against a catalogue built from the warehouse — an invented id becomes a clarifying question, not a query returning nothing.
Deterministic code writes the SQL. Every user-influenced value is a bound parameter. A window function resolves multi-source disagreement by a curated preference recorded in the semantic layer, so a two-country chart can't silently draw four overlapping lines.
Cost per question: about $0.0009.
Full trace, real SQL: groundfact.com/docs/how-a-chart-is-made/
Contact¶
TWAKKA LTD — groundfact.com