Vikram Adithya

Sr Databricks Engineer

US.

About

Highly accomplished Senior Databricks Engineer with 10+ years of expertise in designing, developing, and optimizing scalable data engineering and cloud solutions across healthcare, retail, and financial services. Proven leader in building high-volume ETL/ELT pipelines, Lakehouse architectures, and advanced customer data platforms (AEP, CDP, AJO), driving real-time data activation and customer 360 initiatives. Adept at leveraging Databricks, Spark, Python, SQL, AWS, Azure, and CI/CD-DevOps practices to deliver robust, high-performance data solutions and foster data-driven decision-making.

Work

BCBS

|

Sr Databricks Engineer

, FL, US

→

Summary

Leads the design and implementation of scalable data solutions for enterprise customer data and personalized experiences within Adobe Experience Platform (AEP), Adobe Customer Data Platform (CDP), Real-Time CDP, and Adobe Journey Optimizer (AJO) initiatives.

Highlights

Designed and optimized high-performance data engineering solutions for customer profiles, segmentation, identity resolution, Customer 360, and real-time customer activation across the Adobe Experience Platform ecosystem.

Built enterprise-scale ETL/ELT pipelines using Databricks, Apache Spark, PySpark, Python, SQL, Delta Lake, Azure Data Lake Storage (ADLS), and AWS S3 to ingest, transform, cleanse, enrich, and publish large volumes of structured and semi-structured customer data.

Implemented Delta Lake architecture for reliable data storage and processing, leveraging ACID transactions, schema enforcement, schema evolution, time travel, optimized data layouts, and incremental processing patterns.

Developed and maintained Azure Data Factory pipelines for data ingestion, transformation, scheduling, and dependency management, integrating enterprise source systems with ADLS, Databricks, and downstream platforms.

Integrated streaming and event-driven customer data using Kafka/Event Streaming to support near-real-time ingestion, customer profile updates, segmentation, identity processing, and activation use cases.

Implemented robust data quality, validation, reconciliation, error handling, and monitoring mechanisms across ETL/ELT pipelines, ensuring accuracy, completeness, and reliability of critical customer data.

Utilized Unity Catalog within Databricks to support enterprise data governance, centralized access control, data discovery, permissions management, lineage, and secure management of data assets.

Developed ML feature engineering pipelines using Databricks, Spark/PySpark, Python, and SQL to prepare and manage machine-learning-ready customer features for analytics and predictive use cases.

MasterCard

|

Databricks Engineer

New York, NY, US

→

Summary

Engineered scalable data pipelines and solutions for enterprise customer data initiatives, processing large volumes of customer, transaction, and behavioral data.

Highlights

Developed and maintained end-to-end ETL/ELT pipelines to ingest data from multiple enterprise sources and transform it into reliable datasets for analytics, customer profiles, segmentation, and downstream marketing applications.

Built data processing workflows in Databricks using PySpark and SQL, implementing data cleansing, transformation, aggregation, enrichment, and incremental processing logic for large-scale datasets.

Worked with Adobe Experience Platform (AEP), Adobe Customer Data Platform (CDP), and Real-Time CDP to prepare and integrate customer data required for unified profiles, identity resolution, segmentation, and activation.

Supported Adobe Journey Optimizer (AJO) initiatives by delivering accurate and timely customer datasets used for personalized customer journeys, audience targeting, and customer engagement use cases.

Designed and implemented customer data pipelines supporting Customer 360 capabilities by bringing together customer, purchase, interaction, digital, and behavioral data from multiple enterprise systems.

Developed batch and near-real-time data pipelines using Kafka/Event Streaming to process customer events and make data available for real-time activation and downstream applications.

Optimized Databricks and Spark workloads by reviewing partitioning strategies, joins, transformations, file sizes, and query execution to improve pipeline performance and processing efficiency.

Supported CI/CD implementation using Azure DevOps and Git, automating deployment of Databricks notebooks, PySpark code, SQL scripts, and ADF pipelines across environments.

Wells Fargo

|

Data Engineer

, , US

→

Summary

Designed and implemented scalable data models and ETL pipelines for enterprise reporting and analytics, leveraging Databricks Lakehouse architecture.

Highlights

Designed and implemented scalable conceptual, logical, and physical data models for enterprise reporting and analytics using Databricks Lakehouse architecture.

Developed Bronze, Silver, and Gold data layers using Databricks and Delta Lake to support reliable, reusable, and analytics-ready data products.

Translated business and reporting requirements into scalable data structures, dimensional models, fact tables, dimension tables, and curated reporting datasets, aligning with enterprise data architecture.

Built and optimized ETL/ELT pipelines using AWS Glue, PySpark, Python, and SQL, processing large volumes of structured and semi-structured data for reporting and analytics.

Applied Databricks and Delta Lake modeling best practices, including schema design, partitioning, incremental processing, data validation, and performance optimization, reducing reporting data latency.

Implemented data quality, reconciliation, and validation checks between source systems and curated Databricks datasets to ensure reporting accuracy.

Optimized Spark jobs, SQL queries, and Delta tables to improve data processing performance and reduce reporting data latency.

Supported enterprise reporting teams by creating trusted data marts and curated datasets consumed through Power BI and Tableau, maintaining data lineage and technical documentation.

LTI Mindtree

|

Data Engineer

, , India

→

Summary

Designed and developed scalable data models and ETL pipelines to support enterprise reporting, analytics, and business intelligence requirements.

Highlights

Designed and developed scalable data models and ETL pipelines to support enterprise reporting, analytics, and business intelligence requirements.

Built data pipelines using Apache Spark, PySpark, Hive, Python, and SQL to process large volumes of structured and semi-structured data.

Developed and maintained Bronze, Silver, and Gold data processing layers to organize raw, cleansed, transformed, and reporting-ready datasets.

Implemented dimensional modeling concepts including fact and dimension tables, surrogate keys, business keys, and slowly changing dimensions (SCD).

Integrated data from relational databases, enterprise applications, files, and APIs into centralized Hadoop and cloud-based data platforms.

Developed ETL workflows using AWS Glue and Apache Airflow/Oozie for batch data ingestion, transformation, scheduling, and dependency management.

Migrated and optimized legacy data processing workflows into AWS S3, Redshift, and Databricks environments.

Applied data partitioning, bucketing, file-format optimization, and Spark tuning techniques to improve large-scale data processing performance.

ESM Square Technologies

|

Data Engineer

, , India

→

Summary

Developed and maintained enterprise ETL pipelines and data integration workflows to support reporting, analytics, and business intelligence requirements.

Highlights

Developed and maintained enterprise ETL pipelines and data integration workflows to support reporting, analytics, and business intelligence requirements.

Designed and implemented logical and physical data structures for analytical databases and reporting environments based on business requirements.

Developed ETL workflows using Informatica PowerCenter, SQL, Python, and Shell scripting to extract, transform, validate, and load data from multiple source systems.

Integrated data from Oracle, SQL Server, flat files, and other enterprise applications into centralized data warehouse and Hadoop environments.

Created reusable Informatica mappings, mapplets, workflows, sessions, and transformations for batch data processing.

Implemented data cleansing, standardization, transformation, and business-rule validation to improve overall data quality.

Worked with Hadoop, Hive, and Sqoop to ingest and process large datasets from relational databases into distributed data platforms.

Performed ETL performance tuning by optimizing SQL queries, Informatica workflows, data extraction strategies, and Hadoop processing jobs.

Skills

Data Engineering & Lakehouse

Databricks, Apache Spark, PySpark, Delta Lake, ETL/ELT, Lakehouse Architecture, Data Processing, Batch Processing, Incremental Processing.

Programming & Query Languages

Python, SQL, PL/SQL, Shell Scripting.

Cloud Platforms

Microsoft Azure, AWS, Azure Data Lake Storage (ADLS), AWS S3, AWS Redshift, AWS Glue, AWS EMR, AWS Athena.

Data Integration & Orchestration

Azure Data Factory (ADF), Apache Airflow, Oozie, Informatica PowerCenter, REST APIs, API Integration, Kafka, Event Streaming.

Adobe Experience Platform

Adobe Experience Platform (AEP), Adobe Customer Data Platform (CDP), Adobe Real-Time CDP, Adobe Journey Optimizer (AJO), Customer 360, Customer Identity Resolution, Customer Profile, Audience Segmentation, Real-Time Activation.

Data Modeling & Warehousing

Data Modeling, Dimensional Modeling, Conceptual/Logical/Physical Data Modeling, Star Schema, Fact & Dimension Tables, Slowly Changing Dimensions (SCD), Data Marts, Data Warehousing, Source-to-Target Mapping, Data Lineage.

Databricks & Data Governance

Unity Catalog, Delta Lake, ACID Transactions, Schema Evolution, Schema Enforcement, Data Governance, Access Control, Data Quality, Data Validation, Data Reconciliation.

Data Engineering & Optimization

Spark Optimization, Partitioning, Bucketing, Caching, Efficient Joins, Query Optimization, File-Size Optimization, Performance Tuning, Fault Handling, Error Handling.

Analytics & BI

Power BI, Tableau, Analytical Datasets, Reporting, Business Intelligence.

DevOps & CI/CD

Git, Azure DevOps, Jenkins, CI/CD, Branching, Pull Requests, Code Reviews, Build & Release Management, Automated Deployment.

Big Data Technologies

Hadoop, HDFS, Hive, Sqoop, Apache Spark, Kafka.

Databases

Oracle, SQL Server, PostgreSQL, Snowflake, AWS Redshift.

Machine Learning

ML Feature Engineering, Customer/Behavioral Feature Engineering, Machine-Learning-Ready Data Pipelines.

Containers & Development

Docker, Kubernetes, Linux.

Project & Delivery

JIRA, Agile, Scrum, Production Support, Root Cause Analysis, Technical Documentation.