Over 10+ years of experience designing, developing, and implementing enterprise-scale Data Engineering, Big Data, Cloud Analytics, and Data Warehousing solutions. Extensive experience in architecting scalable Data Lake, Data Warehouse, and Modern Data Platform solutions on AWS and Azure cloud environments. Strong hands-on experience with Apache Spark, PySpark, Spark SQL, Scala, Hadoop Ecosystem, Kafka, Flink, Hive, HBase, and Distributed Data Processing frameworks. Proficient in building enterprise ETL/ELT solutions using AWS Glue, Azure Data Factory, Databricks, EMR, SSIS, and cloud-native data integration services. Experienced in developing large-scale data processing applications using PySpark, Scala, Python, and SQL for business-critical analytics workloads. Strong expertise in designing cloud-based Data Lakes using Amazon S3, Azure Data Lake Storage Gen2, Delta Lake, and Hive-based architectures. Hands-on experience implementing real-time streaming solutions using Apache Kafka, Spark Structured Streaming, Kinesis Firehose, and Apache Flink.
Expertise in designing and optimizing data warehouses using Amazon Redshift, Azure Synapse Analytics, Snowflake, Hive, Teradata, and SQL Server. Proficient in cloud architecture design leveraging AWS services including EMR, S3, Glue, Lambda, Redshift, DynamoDB, Athena, IAM, CloudWatch, SageMaker, Step Functions, and KMS. Experienced with Azure services including Azure Data Factory, Azure Databricks, Azure Synapse Analytics, ADLS Gen2, Azure Key Vault, Azure Monitor, Azure Logic Apps, and AKS. Strong expertise in data ingestion from relational databases, transactional systems, APIs, flat files, vendor feeds, and enterprise applications. Hands-on experience building machine learning data pipelines supporting predictive analytics, model training, deployment, monitoring, and inference workflows. Experienced in workflow orchestration using Apache Airflow, AWS Step Functions, Azure Logic Apps, Oozie, and cloud scheduling frameworks. Strong programming skills in Python, PySpark, Scala, SQL, Shell Scripting, and Bash for large-scale data engineering solutions.
Expertise in designing secure data platforms implementing IAM, Kerberos, encryption, access controls, data lineage, auditing, and compliance standards. Experience working with NoSQL databases including DynamoDB, Cassandra, Cosmos DB, and HBase for scalable data storage solutions. Proficient in developing Spark SQL, Hive SQL, Teradata SQL, and advanced SQL-based transformation and analytical solutions. Hands-on expertise in implementing Infrastructure as Code using Terraform, AWS CloudFormation, and ARM templates. Experienced in designing and implementing CI/CD pipelines using Jenkins, Git, AWS CodePipeline, CodeBuild, Azure DevOps, and CodeDeploy. Strong experience in monitoring and observability solutions using CloudWatch, Elasticsearch, Kibana, Splunk, Azure Monitor, and Log Analytics.
American Express
NY
Designed and implemented a scalable AWS-based Enterprise Data Lake platform to support regulatory reporting, customer analytics, risk management, transaction monitoring, and enterprise-wide data consumption initiatives.
Architected and developed batch and real-time data ingestion frameworks using Amazon EMR, AWS Glue, Amazon Kinesis Firehose, Amazon S3, and Apache Spark for processing high-velocity transactional and market datasets.
Led end-to-end cloud data platform architecture assessments and implemented AWS services including Amazon EMR, Amazon Redshift, Amazon S3, DynamoDB, SageMaker, Lambda, and Step Functions to enable secure and scalable analytics solutions.
Engineered robust data ingestion pipelines integrating data from core transactional systems, payment platforms, market feeds, relational databases, and external vendor sources into centralized cloud-based repositories.
Developed enterprise-grade ETL and ELT frameworks using PySpark, Scala, Spark SQL, AWS Glue, and Hive to support large-scale data transformation and processing requirements.
Built reusable migration accelerators leveraging Spark Data Sources, Hive Metastore, schema evolution strategies, and metadata-driven processing for modernizing legacy data platforms.
Implemented secure data access controls using AWS IAM, Lambda, DynamoDB, bucket policies, and encryption standards to meet stringent governance and compliance requirements.
Established Kerberos-based authentication and authorization mechanisms across Hadoop ecosystems, enabling secure access to HDFS, Hive, Pig, YARN, and MapReduce environments.
Developed high-performance Spark applications using Scala and PySpark to process large-scale structured, semi-structured, and unstructured datasets across distributed environments.
Created optimized Spark SQL solutions utilizing DataFrames, Dataset APIs, partitioning strategies, and caching mechanisms to improve analytical processing efficiency.
Developed advanced Teradata BTEQ scripts and SQL-based transformation processes supporting historical tracking, data reconciliation, deduplication, and master data management requirements.
Managed enterprise data migration and integration initiatives using SQL Server Integration Services (SSIS), DTS packages, and cloud-native data movement services.
Implemented Master Data Management processes and reference data governance standards through data cleansing, transformation, enrichment, and validation methodologies.
Built machine learning data pipelines in Python to support predictive analytics, customer behavior modeling, operational forecasting, and intelligent decision-making use cases.
Automated machine learning workflows using AWS Step Functions and SageMaker for model training, validation, deployment, monitoring, and inference orchestration.
Integrated Apache Airflow with AWS services to orchestrate complex multi-stage workflows across EMR clusters, SageMaker environments, and cloud-based data platforms.
Developed AWS Lambda functions using Boto3 SDK to automate infrastructure operations, governance controls, metadata management, and cloud resource optimization activities.
Optimized Amazon Redshift data warehouses through workload management, query performance tuning, distribution strategies, sort key design, and storage optimization techniques.
Designed and maintained CI/CD pipelines using AWS CodePipeline, CodeBuild, CodeDeploy, Git, and Jenkins to automate deployment and release management processes.
Implemented Infrastructure as Code solutions using Terraform and AWS CloudFormation to provision secure, scalable, and repeatable cloud environments.
Developed centralized observability and monitoring solutions using Elasticsearch, Kibana, CloudWatch, and CloudTrail to support operational intelligence and anomaly detection.
Created governed analytical datasets using Alteryx, advanced SQL, and Tableau to deliver executive dashboards, operational reports, and self-service analytics capabilities.
Established enterprise data governance frameworks encompassing metadata management, lineage tracking, audit controls, encryption using AWS KMS, and regulatory compliance standards.
Led performance tuning initiatives for Spark workloads, EMR clusters, Redshift environments, and distributed data pipelines supporting large-scale enterprise datasets.
Environment: Amazon Web Services, Apache Spark (Scala, PySpark, Spark SQL), Hadoop, Apache Airflow, Elasticsearch, Logstash, Kibana, Teradata (BTEQ), Microsoft SQL Server Integration Services, SQL Server, HBase, Python (Boto3, ML), Alteryx, Tableau, CI/CD, IAM, KMS
Cigna
CT
Built end-to-end data ingestion pipelines using Azure Data Factory to integrate on-premises sources such as MySQL and Cassandra with cloud platforms including Azure Blob Storage, ADLS Gen2, and Azure SQL Database, loading curated data into Azure Synapse Analytics.
Developed batch and real-time data processing pipelines using Azure Databricks with Apache Spark, Spark SQL, PySpark, and Scala for processing claims and patient data feeds.
Architected scalable data lake solutions using Azure Data Lake Storage Gen2 following medallion architecture principles for structured and semi-structured healthcare datasets.
Migrated large-scale data from Hadoop and HDFS environments to Azure Databricks and Azure Synapse Analytics while ensuring data integrity and governance.
Created and administered Databricks clusters with auto-scaling configurations and implemented Spark performance tuning strategies.
Mounted Azure Blob Storage and ADLS Gen2 within Databricks for distributed processing and data curation.
Built Spark Streaming applications consuming real-time event data from Kafka using both stateless and stateful transformations.
Implemented ETL and ELT frameworks using Azure Data Factory, Databricks, Hive, and Snowflake SnowSQL to process structured and semi-structured data formats including JSON, Avro, Parquet, ORC, XML, and CSV.
Developed Hive data warehouse solutions implementing partitioning and bucketing strategies to optimize query performance.
Engineered automated workflows using Azure Logic Apps and orchestrated containerized workloads using Azure Kubernetes Service for scalable data processing.
Deployed and managed Spark applications in Kubernetes environments to support distributed analytics.
Monitored Spark and Hadoop clusters using Azure Log Analytics, Azure Monitor, and Ambari Web UI to ensure operational stability.
Transitioned log storage from Cassandra to Azure Synapse dedicated SQL pool to enhance reporting and analytics capabilities.
Implemented secure data pipelines using Azure Key Vault for secrets management and deployed solutions across environments using ARM templates and Azure DevOps CI/CD practices.
Developed reusable Databricks notebooks and PySpark libraries to standardize data transformation, cleansing, and enrichment processes.
Worked extensively with Azure Cosmos DB using SQL API and Mongo API for scalable storage of -related semi-structured data.
Developed MapReduce jobs, Pig scripts, Hive queries, and custom UDFs for large-scale data transformations within Hadoop ecosystems.
Utilized DataStax Spark Connector to integrate Cassandra with Spark-based analytics platforms.
Designed and maintained Oozie workflows for scheduling and managing complex Hadoop job dependencies.
Worked in hybrid cloud environments leveraging AWS Glue with PySpark dynamic frames, crawlers, and workflow scheduling for cross-platform data integration.
Developed interactive dashboards and enterprise reports using SSRS and Power BI, including drill-down, parameterized, matrix, and executive reporting dashboards for stakeholders.
Environment: Azure, Apache Spark (Spark SQL, PySpark, Scala), Delta Lake, Apache Kafka, Apache Hive, Apache Hadoop (HDFS, MapReduce, Pig, Oozie), Apache Cassandra, Azure Cosmos DB (SQL & Mongo API), Snowflake, Power BI, SQL Server Reporting Services, Azure DevOps, Azure Key Vault; MySQL, JSON
Fieh Third Bank
PA
Designed and developed scalable batch and real-time data processing applications using Apache Spark, PySpark, Scala, Kafka, Hive, and Hadoop ecosystem technologies to process large volumes of transactional and operational data.
Built robust data ingestion frameworks to collect and integrate structured, semi-structured, and unstructured data from payment platforms, CRM systems, operational databases, flat files, and third-party service APIs.
Developed Spark SQL and PySpark-based ETL pipelines to cleanse, validate, standardize, enrich, and transform raw data into analytics-ready datasets for reporting and advanced analytical use cases.
Implemented Apache Kafka-based streaming solutions to capture, process, and distribute event-driven transactional data across enterprise data platforms with low-latency processing requirements.
Designed and developed Apache Flink streaming applications to consume Kafka events, perform real-time transformations, enrich incoming datasets, and publish processed records for downstream analytical systems.
Utilized Apache Avro schemas for data serialization, schema evolution, and efficient data exchange across distributed processing environments and streaming platforms.
Developed Spark Structured Streaming applications to process near real-time transaction data, perform anomaly identification, and persist curated datasets into HBase and Amazon S3 repositories.
Created reusable shell scripts, automation utilities, and framework components for Hive, Sqoop, Flink, Spark, and Pig workloads to streamline data engineering development activities.
Developed Hive external and managed tables, optimized partitioning strategies, and implemented efficient querying techniques to support large-scale analytical workloads.
Performed data extraction from relational databases using Sqoop and integrated high-volume datasets into distributed Hadoop environments for enterprise analytics consumption.
Assisted in administration, monitoring, troubleshooting, and operational support of Hadoop, Spark, Hive, Kafka, Splunk, and Tableau environments.
Designed and maintained Tableau dashboards and analytical reports that enabled business users to monitor operational trends, transaction activity, customer behavior, and performance metrics.
Developed REST API integrations to exchange data between enterprise applications, external systems, and cloud-based platforms while ensuring secure and reliable data movement.
Built scalable cloud-based data ingestion and processing solutions using Amazon EMR, Apache Spark, Python, and Amazon S3 for enterprise data lake initiatives.
Worked extensively with AWS services including EMR, S3, Glue, Lambda, Redshift, Athena, IAM, EC2, VPC, CloudWatch, and SNS to support cloud-native data engineering solutions.
Developed AWS Glue ETL jobs using PySpark and Spark SQL to transform and load enterprise datasets into Redshift and other analytical repositories.
Optimized Spark and PySpark application performance through effective partition management, join optimization techniques, memory tuning, caching strategies, and resource allocation improvements.
Assisted in building CI/CD pipelines using Git, Jenkins, AWS CodePipeline, and AWS CodeBuild to automate application deployment, testing, and release management processes.
Environment: Apache Spark, PySpark, Scala, Apache Kafka, Apache Flink, Hadoop, HDFS, Hive, Sqoop, Pig, Spark SQL, Spark Structured Streaming, HBase, Apache Avro, Splunk, Tableau, Python, Shell Scripting, REST APIs, AWS, Git, Jenkins, CI/CD, Data Lake, ETL, Data Warehousing, Real-Time Streaming, Data Governance, Agile, SDLC
Teradata
India
Migrated enterprise databases from SQLite to MySQL and PostgreSQL while maintaining data integrity, schema consistency, and seamless business continuity throughout the migration lifecycle.
Developed and optimized complex SQL queries, joins, views, stored procedures, and database objects to support reporting, analytics, and operational data requirements.
Performed data mapping, schema analysis, and data transformation activities to ensure consistency between source and target systems during migration projects.
Built scalable data processing solutions using Python to handle structured, semi-structured, and transactional datasets from diverse sources.
Developed automated data ingestion and conversion frameworks using Python and Bash scripting to streamline recurring integration processes.
Created reusable Python modules for data cleansing, validation, normalization, and transformation to improve overall data quality and consistency.
Implemented database connectivity solutions using Python, SQLAlchemy, and database connectors for efficient data access and integration.
Developed RESTful APIs using Python and Django to enable secure data exchange and integration between internal and external systems.
Integrated messaging platforms using AMQP and RabbitMQ to support asynchronous data processing and event-driven data workflows.
Built data pipelines for processing, aggregating, and transforming large datasets to support reporting and operational business functions.
Utilized Python libraries including Pandas, NumPy, Itertools, Functools, Glob, Random, and Math for advanced data manipulation and transformation tasks.
Developed data visualizations using Matplotlib, Seaborn, Highcharts, HTML, CSS, and JavaScript-based components for analytical reporting.
Generated operational and analytical reports using SQL Server Reporting Services (SSRS), including tabular, matrix, chart, and parameterized reports.
Customized SSRS reporting solutions using URL access configurations to support dynamic report generation and distribution requirements.
Performed root cause analysis, debugging, and issue resolution for database, ETL, and integration-related challenges.
Environment: Python, Django, SQLAlchemy, MySQL, PostgreSQL, SQLite, SQL Server, RabbitMQ, Git, Postman, SSRS, HTML, CSS, jQuery, Matplotlib, Seaborn, RESTful API, ETL
Hadoop, HDFS, YARN, MapReduce, Hive, HBase, Sqoop, Pig, Flume, NiFi, Oozie, Kafka, Zookeeper, Apache Spark (Scala, PySpark, Spark SQL, DataFrame API, MLlib, Spark Streaming), Flink, Mahout
AWS (S3, EMR, Glue, Lambda, Redshift, Athena, DynamoDB, SageMaker, CloudWatch, SNS, Step Functions, CodePipeline, CodeBuild, KMS, IAM), Azure (Data Factory, Databricks, ADLS Gen2, Synapse Analytics, Logic Apps, Key Vault, Azure DevOps, ARM templates)
Oracle, SQL Server, MySQL, PostgreSQL, Teradata, MongoDB, Cassandra, DynamoDB, Cosmos DB (SQL API & Mongo API)
Star Schema, Snowflake Schema, Dimensional Modeling, ER Modeling, Transactional Modeling, Slowly Changing Dimensions (SCD Type 2)
Spark, Hive, Sqoop, Flume, Pig, SSIS, Azure Data Factory, DataBricks, ELT frameworks, Data Pipelines (Batch & Real-Time), Delta Lake, Spark Structured Streaming
Python, PySpark, Scala, Java, C, C++, Shell Script (Bash), Perl, SQL
Django REST framework, MVC, Hortonworks
Spark MLlib, Python ML libraries, Linear & Logistic Regression, Decision Trees, Random Forest, NLP, Clustering, Associative Rules
Tableau, Power BI, SSRS, Matplotlib, Seaborn, ggplot2, Highcharts
Apache Airflow, Splunk, ELK Stack (Elasticsearch, Logstash, Kibana), CloudWatch, Azure Monitor, Ambari
Git, SVN, GitHub, Jenkins, AWS CodePipeline, CodeBuild, Terraform, Ansible, Docker
RESTful APIs, JSON/XML, RabbitMQ