← All projects

PUBLIC ENGINEERING CASE STUDY

DataNepal

A public data product with canonical geography, required provenance, and a clear boundary between ingestion, transformation, and publication.

The problem

Nepal’s public statistics are spread across government sites. Sources use different formats, geography codes, and update cadences. A question spanning population and geography can become a reconciliation exercise.

DataNepal conforms datasets to the same province, district, and local-unit model while preserving where their numbers came from. This public project provides inspectable implementation evidence alongside my enterprise experience.

The architecture

GitHub Actions runs ingestion, transformation, and tests, then exports published files. Those files are versioned in the repository so the web deployment can build with Node without rerunning the Python pipeline.

The boundaries are explicit: ingestion handles source behavior; dbt handles the analytical model; export enforces the publication contract.

Design decisions

Canonical geography is the join key

OCHA P-codes anchor a shared geographic spine. Other sources reach it through crosswalk tables, tested for one-to-one mappings and complete coverage. This avoids relying on inconsistent English and Nepali name matching.

Provenance is part of the schema

Each published table needs a catalog entry with its source, licence, vintage, and caveats. The exporter refuses a table without one. Downloads retain the context needed to interpret their data.

Aggregates define the publication boundary

The platform publishes aggregate statistics. Dataset metadata explicitly asserts that published tables contain no personal data. Sources that cannot satisfy the project’s publication constraints are excluded.

The serving tradeoff

The documented data footprint fits Parquet and JSON delivered through a CDN, with DuckDB-WASM handling analytical queries in the browser. This keeps an always-running query service out of the operating model.

The product serves published snapshots. Refreshes pass through ingestion, validation, and export. Data volume and query requirements should be revisited before extending this approach to larger or continuously changing workloads.

This tradeoff is an engineering interpretation of the documented architecture, not a measured performance or cost claim.

Inspect the evidence

The live site covers 7 provinces, 77 districts, and 753 local units. These describe geographic coverage, not adoption or business impact.

Based on the public project documentation and live site reviewed October 2, 2026.

← All projects