ENTERPRISE DATA PLATFORMS

Roshan
Dhamala.

Senior Data Engineer

From platform foundations
to production data & AI.

I help take enterprise lakehouses from proof of concept to production—setting integration standards, improving Spark performance, and building governed data and AI workflows.

Veeva Systems · 2022–present · Senior since 2025

ENGINEERING IMPACT & PLATFORM SCALE

Selected outcomes & scale.

More context

≈35%

Lower processing cost

Reduction in Spark-related processing cost.

≈39.6M

Historical records

Historical records processed in a major ingestion initiative.

100+ TB

Lakehouse volume

Platform scale: the production Databricks lakehouse I helped evolve.

120+

Datasets

Platform scale: datasets across the enterprise lakehouse.

Cost reduction and records processed are approximate professional outcomes, separate from the ongoing CRM migration. Volume and dataset counts describe platform scale, not sole individual accomplishment.

SELECTED ENTERPRISE WORK

Ownership across the
data platform.

All work

ENTERPRISE LAKEHOUSE

From POC
to production.

A shared Databricks foundation
for operational and analytical data.

PLATFORM ARCHITECTURE & ENGINEERING

LAKEHOUSE ARCHITECTURE

Engineering standards across the platform.

I helped evolve a Databricks lakehouse from proof of concept to production, defining ingestion, modeling, and governance standards used by engineering and analytics teams.

  • Batch and near-real-time integration reference patterns
  • Bronze/silver/gold modeling with Delta Lake and dbt
  • Unity Catalog, data contracts, lineage, and access control
  • Spark performance work, architecture reviews, and mentoring
Read the platform overview

APPLIED DATA + AI

Documents → usable data

LLM/RAG workflows for document processing, structured extraction, metadata enrichment, and vector search over governed enterprise data.

Engineering scope

ONGOING · CRM MIGRATION

Cross-system data continuity.

Technical ownership of data engineering across source-to-target mapping, reconciliation, integration, reporting, and permissions.

Migration scope

PUBLIC ENGINEERING PROJECT

DataNepal.

Inspectable code. Documented decisions.

A public data product that brings fragmented sources into a shared geographic model, with provenance and a static publishing boundary.

dltDuckDBdbtParquet

What the implementation demonstrates

  • Canonical identifiers and tested geographic crosswalks
  • Source, licence, vintage, and caveats enforced at export
  • Parquet / JSON delivery with browser-side queries
Inspect the repository

TECHNICAL PERSPECTIVE

Decisions, explained.

Writing

CURRENT ROLE

Senior Data Engineer

Veeva Systems · January 2025–Present

At Veeva since January 2022; promoted to Senior Data Engineer in January 2025. Lead architecture reviews and mentor engineers on Spark optimization, modeling, and production patterns.

View resume

CONTACT

Let’s talk data platforms.

Connect with me about enterprise data platforms, technical leadership, or applied AI.

Connect on LinkedIn