About Me
01 / 07
โš™๏ธ Expertise

Highly skilled Data Architect specializing in building scalable, metric-driven search engines and AI-orchestrated architectures. Proven track record in leading engineering teams to solve complex natural language and enterprise data challenges.

๐ŸŽฏ Core Focus

Dedicated to bridging the gap between natural language and structured data (NL2SQL), optimizing semantic precision through advanced retrieval algorithms, and deploying high-performance data frameworks like Google Cortex.

Engineering Leadership

Currently leading a high-performing team of 3 engineers to design and integrate a scalable, headless BI system.

Focused on precision natural language-to-SQL conversions, transforming how analytical stakeholders interact with complex data repositories.

Live Telemetry Metrics
NL2SQL SLA 99.82%
Avg Latency 182ms
Query Accuracy 95.4%
๐Ÿ“ฆ Lead Architect

Driving Looker integration using the Multi-Agent Google ADK engine.

๐Ÿค– Semantic Roles

Configuration of orchestrator and semantic agent roles for intelligent discovery.

๐Ÿ”ฑ Engine Mastery

Managing complex agentic workflows within enterprise environments.

User Orch Agent Tool
๐Ÿ—„๏ธ
Metadata Architecture

Architected robust ingestion workflows from Looker to PostgreSQL storage.

๐Ÿ› ๏ธ
Context Management

Implementation of custom ADK context management for high data fidelity.

โšก
System Performance

Optimized storage layers for rapid analytical query response times.

Agentic Protocol Stack
ANP Discovery & Identity
A2A Agent Coordination
MCP Agent-to-Tool Access
RPC JSON-RPC 2.0 Format
RRF Reciprocal Rank Fusion

Optimizing Semantic Precision

Developed advanced retrieval algorithms combining Reciprocal Rank Fusion (RRF) and Full-Text Search (FTS). Utilizing knowledge recipes to ensure top-tier accuracy.

RRF Merging Simulator
FTS (Lexical) Semantic (Vector)
Query Match Rank
FTS: #12
Vector: #5
RRF: #1
โ˜๏ธ
Strategic Deployment

Co-led the deployment of the Google Cortex framework for enterprise reporting.

๐Ÿ“ˆ
BigQuery Native

Empowering cross-functional SAP reports natively within BigQuery.

โšก
Operational Impact

Streamlining SAP data pipelines to provide real-time insights.

Data Foundation Pipeline
๐Ÿ—„๏ธ SAP ERP
โ†’
Cortex BigQuery
โ†’
๐Ÿ“Š Looker BI
๐Ÿ“ Model Context Protocol

MCP Server Integration

Designing structured multi-agent development workflows via Model Context Protocol (MCP), exposing local systems as secure tools for LLMs.

๐Ÿงฌ Repository Scale

Context Management

Enforcing workflows inside expanding repository-level contexts for seamless scalability, keeping memory usage bounded while retaining full search.

Host Agent MCP Servers
filesystem websearch postgres
Hover over tool nodes to request execution context.
AI & Data Engineering Architect

Jatin Aggarwal

Data Platform & Generative AI Workflows

Data Platform & AI Architect with nearly 8 years of experience pioneering the convergence of enterprise data engineering and generative AI. Specialist in building self-serve data ecosystems, headless BI, and advanced semantic layers that transform raw enterprise data into LLM-ready intelligence. Proven track record leading engineering teams to architect multi-agent AI frameworks (Google ADK, Vertex Agent Engine), production-grade vector search systems, and high-performance cloud data platforms. Expert at building the orchestration, ingestion, and advanced retrieval algorithms (RRF, hybrid lexical/FTS) required to deliver precise, secure, and autonomous natural language-to-data experiences at enterprise scale.

Interactive Conversation Starters:
8
Years Experience
~60%
Compute Cost Reduction
~85%
decrease in discovery time as stated by daily active user
~50%
Processing Improvement
50+
Stakeholders Served
[SYSTEM] Humor Telemetry Console ONLINE
[QUOTE] dbt compile: turning SELECT * into a million-dollar cloud bill.
System Console

AI Agent Live Flow Simulator

ADK Agent Execution Trace
VISUAL PATHWAY
1. INGESTION ENGINE (Looker Metadata Sync)
Metadata Fetch
Hash Verify
Create Recipe
Embeddings
2. RETRIEVAL & MULTI-AGENT EXECUTION (ADK)
Natural Language query received NL User Query Classify intent and select tool routing ORCH Orchestrator Fetches web results for external queries SRCH Web Search Reciprocal Rank Fusion, Lexical check and Recipe lookup SUB Subagent (RRF) Asks user to clarify when score is below threshold CLRF Clarify BigQuery / Looker SDK query validation BQ Tools Exec Data rendered as intention OUT Data Plot
Query Controller
ACTIVE
Agent Telemetry & Eval
OTEL SPANS
Eval Accuracy 1.0
Confidence Score -
SLA Latency -
OpenTelemetry Spans (Span List):
Execute a query trace to capture active trace telemetry spans.
Execution Console & Output Visualizer
Agent Thinking Logs:
[SYSTEM] Ready for telemetry. Pick a query scenario and click 'Execute Trace Run'.
Dynamic Output Canvas: (Real-time Headless BI & Agent Execution Outputs)

Execution results will be plotted here.

Workspace Configurations

Code & Semantic Explorer

jatin@data-platform:~/portfolio
INDEXED
select a file to view
TEXT
/* Select a file from the workspace tree to view source code & system logs */
Registry

Projects & Articles

project

Natural Language Data Agent

A domain-agnostic AI agent that lets analysts query data, helps engineers trace pipelines, and gives EMs visibility into team work, all from plain-English questions backed by versioned data contracts.

AI AgentsBigQueryNLPData EngineeringLLM
~20-30 Daily Active User
~50-100 Daily Active Runs
Explore Case Study
article

Multi-Agent Development Framework

Designed a multi-agent development framework in Cursor, improving development speed and consistency across design, build, test, and documentation.

CursorAI AgentsMCPProductivity
20-30% velocity boost
Explore Case Study
project

AI-Powered Analytics Assistant

Built an AI-powered analytics assistant that lets business users pull real-time data independently, reducing repeated data-fetch requests to the data engineering team.

AIBigQueryPython
10โ€“100 daily queries
5โ€“10 business users
Explore Case Study
project

Near Real-Time Data Ingestion System

Built a near real-time data ingestion system, reducing compute cost by ~60% and simplifying architecture.

Pub/SubCloud FunctionsBigQueryTerraform
~60% cost reduction
NRT latency
Explore Case Study
project

Large-Scale Traceability System

Developed a large-scale traceability system supporting high-volume operational data, improving visibility for business decisions.

AirflowdbtBigQuery
1M+/day volume
Explore Case Study

Jatin Aggarwal

Senior Data Engineer & AI Architect ยท Data Platform & Generative AI Workflows

Email: jatinagg1307@gmail.com Phone: +91 9557377987 Location: Bangalore, India LinkedIn: linkedin.com/in/jatin-aggarwal-451041116

Professional Summary

Data Platform & AI Architect with nearly 8 years of experience pioneering the convergence of enterprise data engineering and generative AI. Specialist in building self-serve data ecosystems, headless BI, and advanced semantic layers that transform raw enterprise data into LLM-ready intelligence. Proven track record leading engineering teams to architect multi-agent AI frameworks (Google ADK, Vertex Agent Engine), production-grade vector search systems, and high-performance cloud data platforms. Expert at building the orchestration, ingestion, and advanced retrieval algorithms (RRF, hybrid lexical/FTS) required to deliver precise, secure, and autonomous natural language-to-data experiences at enterprise scale.

Core Skills

Data Platforms Google BigQuery, Databricks, dbt, Cloud Storage, Data Mesh, Terraform
Data Engineering Multi Agent Google ADK, Vertex Agent Engine, Vector Databases (Vertex AI Vector Search, Qdrant), Semantic Engines (Cube Core, Looker, Looker SDK, LookML), Semantic Model Designing, Data Contracts, Service Level Agreements (SLAs), Source-to-Target Lineage, Schema Design, Data Quality
Distributed Systems Cloud Composer (Apache Airflow), Google Cloud Dataflow (Apache Beam), Google Cloud Pub/Sub, Spark
AI & Productivity Context Management, MCP Servers, Reciprocal Rank Fusion (RRF), Hybrid Lexical/FTS
Languages Python, SQL, Java (Basics), LLM-driven development (Cursor, Antigravity, Claude Code, Glean)

Professional Experience

Data Architect - Data Engineer Dec 2024 - Jun 2026
Wayfair ยท Bangalore
  • Leading a team of 3 engineers to design and integrate a scalable, metric-driven semantic layer within LLMs to enable headless BI and precise natural language-to-SQL conversions
  • Serving as Lead Architect on Looker integration using the Multi-Agent Google ADK framework hosted on Vertex Agent Engine, configuring orchestrator and semantic agent roles
  • Architected ingestion workflows from Looker to PostgreSQL metadata storage, implementing hash-based embeddings and customized ADK context management
  • Developed advanced retrieval algorithms combining Reciprocal Rank Fusion (RRF) with full-text search (FTS), lexical FTS, and knowledge recipes to optimize semantic precision
  • Co-led deployment of the Google Cortex framework, empowering cross-functional analytical teams to query complex SAP operational reports natively within BigQuery
  • Designed and enforced structured multi-agent development workflows inside Cursor, utilizing Glean agents and expanding repository-level context via Model Context Protocol (MCP) servers
  • Created an AI-driven automated analytics assistant utilizing Python, minimizing repetitive ad-hoc query fulfillment cycles for analytical stakeholders
  • Spearheaded end-to-end migration of high-priority enterprise datasets to the central GCP platform, achieving 50% faster processing via highly optimized dbt models
  • Accelerated enterprise data mesh adoption across several organizational domains, substantially increasing cross-domain asset reuse and discoverability
  • Architected a real-time streaming data ingestion pipeline from Google Cloud Firestore to BigQuery, delivering a 60% reduction in platform compute costs
  • Implemented a dbt-based micro-batch pipeline to process information from heterogeneous sources, providing unified traceability of carton lifecycle events
  • Mentored junior engineers and led technical onboarding on data mesh principles, production dbt modeling, and optimal GCP service configurations
Senior Data Engineer (Contract) Nov 2023 - Dec 2024
Wayfair ยท Remote
  • Constructed and scaled foundational data pipelines with Python, BigQuery, and dbt to service over 50 downstream enterprise stakeholders, analytical dashboards, and key business decision matrices
  • Enhanced pipeline performance by 30-35% and lowered cloud operational spend by 20-25% via targeted SQL query optimizations and advanced BigQuery feature utilization
  • Created highly modular data transformation architectures with dbt that minimized manual modeling steps and accelerated iteration velocity
  • Defined and executed strict data contracts and service level agreements (SLAs) utilizing GCP tools, leading to a substantial decrease in downstream data quality incidents
Cloud Data Engineer Aug 2022 - Nov 2023
Impetus Technologies ยท India
  • Modernized and migrated over 40 legacy relational warehouse workflows to modern cloud lakehouse frameworks leveraging Python, Databricks on GCP and AWS with no data loss
  • Architected robust batch processing pipelines processing multi-terabyte weekly data scales with Spark hosted on scalable cloud infrastructure
  • Deployed automated monitoring and proactive alerting frameworks across data assets, decreasing manual operational monitoring workloads by 60%
Data Engineer Oct 2021 - Aug 2022
Futurense Technologies (Rakuten India) ยท India
  • Developed production-grade batch and streaming pipelines via Python, BigQuery, and dbt for customer and market data domains, empowering ML and analytics workloads
  • Enabled near-real-time analytical capabilities with sub-second response times by integrating high-throughput Python microservices with BigQuery architectures
  • Improved downstream data asset accessibility by designing and refactoring dimensional data models using dbt
Data Engineer / QA Automation Engineer Jul 2018 - Oct 2021
Wipro Limited ยท India
  • Built custom automated regression and system verification tools using Python, diminishing manual efforts by 40-50%
  • Developed API testing frameworks and transaction validation scripts in Python to secure enterprise integration reliability
  • Transformed complex functional requirements into lean, generic, and highly reusable technical scripts using Python

Education & Certifications

B.Tech, Computer Science Aug 2014 - May 2018
GLA University
  • Applied Data Science with Python
  • Databricks Certified Associate Developer