RESUME

Resume.

Download my resume, or read the career summary below.

Roshan Dhamala

Senior Data Engineer

LinkedIn · roshandhamala.com · github.com/RDhamala

Professional focus

I design and evolve enterprise data platforms—from lakehouse architecture and integration standards to governed analytics and applied AI. At Veeva, my work has helped take a Databricks platform from proof of concept to production.

Experience

Senior Data Engineer · Veeva Systems
January 2025–Present

  • From POC to production: Technical architecture for a Databricks lakehouse with 100+ TB and 120+ datasets. Defined ingestion, modeling, and governance standards used by engineering and analytics teams.
  • Integration reference design: End-to-end architecture for batch and near-real-time integrations across SaaS, REST APIs, AWS Kinesis, and SFTP, using a bronze/silver/gold medallion design.
  • Performance & trust: Photon, incremental processing, partition pruning, and Spark tuning; governance and reliability patterns using Unity Catalog, data contracts, lineage, RBAC, and automated data quality.
  • Applied AI & technical leadership: LLM/RAG workflows for document processing, structured extraction, metadata enrichment, and vector search. Lead architecture reviews and mentor engineers on Spark, data modeling, and production design.

Earlier experience

Data Engineer · Veeva Systems
January 2022–January 2025

  • Built and operated production batch and streaming pipelines with Databricks, Spark, Delta Lake, AWS Kinesis, and REST APIs.
  • Standardized orchestration, CI/CD, testing, monitoring, and alerting with Apache Airflow, Databricks Jobs, and Git. Developed ETL/ELT models in PySpark, SQL, and dbt.

Software Developer Intern · Texas Tech University Health Sciences Center
January 2021–December 2021

  • Designed Snowflake data models and incremental ingestion with Apache Airflow and Snowflake Tasks for near-real-time reporting.

Software Engineering Intern · Discover Financial Services
June 2021–August 2021

  • Developed ETL workflows for financial transaction data, including profiling, validation, and anomaly detection to support SOX-compliant analytics.

Current initiative: enterprise CRM migration

Technical owner for data engineering. Scope includes migration into the new CRM; source-to-target mapping; reconciliation and data quality; integration architecture and ingestion; operational and analytical reporting; downstream modeling; permissions/governance; and business and technical stakeholder coordination.

Selected professional impact

  • Approximately 35% — Reduction in Spark-related processing cost.
  • Approximately 39.6M — Historical records processed in a major ingestion initiative.

Approximate figures across professional work; separate from the ongoing CRM migration.

Technical scope

Lakehouse & analytics: Databricks, Apache Spark, Delta Lake, Unity Catalog, Snowflake, dbt Cloud

Languages & integration: Python, SQL, PySpark, Structured Streaming, REST APIs, AWS Kinesis, SFTP

Production engineering: Apache Airflow, Databricks Jobs, Git, CI/CD, Data contracts, Lineage, RBAC, Data quality

Applied AI: LLMs, RAG, Document processing, Structured extraction, Embeddings, Vector search, AI agents

Public engineering work

DataNepal: canonical geographic modeling, provenance enforced at export, and a static data publishing architecture using dlt, DuckDB, dbt, Parquet/JSON, and DuckDB-WASM.

datanepal.org · Source repository

Education

M.S. Software Engineering · Texas Tech University
December 2021

B.S. Computer Science · Texas Tech University
December 2020

Credentials listed on my resume

  • Databricks Certified Data Engineer Associate
  • AWS Certified Cloud Practitioner