Jatin Aggarwal
Data Platform & Generative AI Workflows
Data Platform & AI Architect with nearly 8 years of experience pioneering the convergence of enterprise data engineering and generative AI. Specialist in building self-serve data ecosystems, headless BI, and advanced semantic layers that transform raw enterprise data into LLM-ready intelligence. Proven track record leading engineering teams to architect multi-agent AI frameworks (Google ADK, Vertex Agent Engine), production-grade vector search systems, and high-performance cloud data platforms. Expert at building the orchestration, ingestion, and advanced retrieval algorithms (RRF, hybrid lexical/FTS) required to deliver precise, secure, and autonomous natural language-to-data experiences at enterprise scale.
AI Agent Live Flow Simulator
Execution results will be plotted here.
Code & Semantic Explorer
/* Select a file from the workspace tree to view source code & system logs */ Projects & Articles
Natural Language Data Agent
A domain-agnostic AI agent that lets analysts query data, helps engineers trace pipelines, and gives EMs visibility into team work, all from plain-English questions backed by versioned data contracts.
Multi-Agent Development Framework
Designed a multi-agent development framework in Cursor, improving development speed and consistency across design, build, test, and documentation.
AI-Powered Analytics Assistant
Built an AI-powered analytics assistant that lets business users pull real-time data independently, reducing repeated data-fetch requests to the data engineering team.
Near Real-Time Data Ingestion System
Built a near real-time data ingestion system, reducing compute cost by ~60% and simplifying architecture.
Large-Scale Traceability System
Developed a large-scale traceability system supporting high-volume operational data, improving visibility for business decisions.
Jatin Aggarwal
Senior Data Engineer & AI Architect ยท Data Platform & Generative AI Workflows
Professional Summary
Data Platform & AI Architect with nearly 8 years of experience pioneering the convergence of enterprise data engineering and generative AI. Specialist in building self-serve data ecosystems, headless BI, and advanced semantic layers that transform raw enterprise data into LLM-ready intelligence. Proven track record leading engineering teams to architect multi-agent AI frameworks (Google ADK, Vertex Agent Engine), production-grade vector search systems, and high-performance cloud data platforms. Expert at building the orchestration, ingestion, and advanced retrieval algorithms (RRF, hybrid lexical/FTS) required to deliver precise, secure, and autonomous natural language-to-data experiences at enterprise scale.
Core Skills
Professional Experience
- Leading a team of 3 engineers to design and integrate a scalable, metric-driven semantic layer within LLMs to enable headless BI and precise natural language-to-SQL conversions
- Serving as Lead Architect on Looker integration using the Multi-Agent Google ADK framework hosted on Vertex Agent Engine, configuring orchestrator and semantic agent roles
- Architected ingestion workflows from Looker to PostgreSQL metadata storage, implementing hash-based embeddings and customized ADK context management
- Developed advanced retrieval algorithms combining Reciprocal Rank Fusion (RRF) with full-text search (FTS), lexical FTS, and knowledge recipes to optimize semantic precision
- Co-led deployment of the Google Cortex framework, empowering cross-functional analytical teams to query complex SAP operational reports natively within BigQuery
- Designed and enforced structured multi-agent development workflows inside Cursor, utilizing Glean agents and expanding repository-level context via Model Context Protocol (MCP) servers
- Created an AI-driven automated analytics assistant utilizing Python, minimizing repetitive ad-hoc query fulfillment cycles for analytical stakeholders
- Spearheaded end-to-end migration of high-priority enterprise datasets to the central GCP platform, achieving 50% faster processing via highly optimized dbt models
- Accelerated enterprise data mesh adoption across several organizational domains, substantially increasing cross-domain asset reuse and discoverability
- Architected a real-time streaming data ingestion pipeline from Google Cloud Firestore to BigQuery, delivering a 60% reduction in platform compute costs
- Implemented a dbt-based micro-batch pipeline to process information from heterogeneous sources, providing unified traceability of carton lifecycle events
- Mentored junior engineers and led technical onboarding on data mesh principles, production dbt modeling, and optimal GCP service configurations
- Constructed and scaled foundational data pipelines with Python, BigQuery, and dbt to service over 50 downstream enterprise stakeholders, analytical dashboards, and key business decision matrices
- Enhanced pipeline performance by 30-35% and lowered cloud operational spend by 20-25% via targeted SQL query optimizations and advanced BigQuery feature utilization
- Created highly modular data transformation architectures with dbt that minimized manual modeling steps and accelerated iteration velocity
- Defined and executed strict data contracts and service level agreements (SLAs) utilizing GCP tools, leading to a substantial decrease in downstream data quality incidents
- Modernized and migrated over 40 legacy relational warehouse workflows to modern cloud lakehouse frameworks leveraging Python, Databricks on GCP and AWS with no data loss
- Architected robust batch processing pipelines processing multi-terabyte weekly data scales with Spark hosted on scalable cloud infrastructure
- Deployed automated monitoring and proactive alerting frameworks across data assets, decreasing manual operational monitoring workloads by 60%
- Developed production-grade batch and streaming pipelines via Python, BigQuery, and dbt for customer and market data domains, empowering ML and analytics workloads
- Enabled near-real-time analytical capabilities with sub-second response times by integrating high-throughput Python microservices with BigQuery architectures
- Improved downstream data asset accessibility by designing and refactoring dimensional data models using dbt
- Built custom automated regression and system verification tools using Python, diminishing manual efforts by 40-50%
- Developed API testing frameworks and transaction validation scripts in Python to secure enterprise integration reliability
- Transformed complex functional requirements into lean, generic, and highly reusable technical scripts using Python
Education & Certifications
- Applied Data Science with Python
- Databricks Certified Associate Developer