Abdul Baseer Shaik
00 — LOADING PORTFOLIO
Available for AWS data engineering roles

Abdul Baseer Shaik.

AWS Data Engineer building cloud data pipelines, enterprise data warehouses, and production-ready ETL/ELT systems across AWS, Spark, and Snowflake.

Snapshot/ 2026
0+
Years in data engineering
0
Industry domains
0+
Core AWS data services
M.S.
Information Technology
Scroll to explore
AWS GlueAmazon S3Amazon EMRRedshift LambdaStep FunctionsSnowflakePySpark Spark SQLScalaHadoopKafka AirflowInformaticaData Warehousing AWS GlueAmazon S3Amazon EMRRedshift LambdaStep FunctionsSnowflakePySpark Spark SQLScalaHadoopKafka AirflowInformaticaData Warehousing
01

Profile

I build production data systems that move, transform, validate, and serve enterprise data across AWS, Hadoop, Spark, Snowflake, and relational platforms.

AWS Data Engineer with 12+ years of experience designing, modernizing, and supporting cloud data platforms, big data workloads, ETL/ELT pipelines, and enterprise data warehouses.

My work spans AWS Glue, Amazon S3, EMR, Redshift, Lambda, Step Functions, Snowflake, PySpark, Spark SQL, Scala, Hadoop, Hive, Kafka, Informatica, Airflow, and CI/CD pipelines for batch and near-real-time data delivery.

My delivery focus includes source-to-target controls, data quality, performance tuning, dimensional modeling, orchestration, security, and dependable production support across six industry domains.

At a glance

Based inUnited States
DisciplineAWS Data Engineering
Experience12+ years
EducationM.S. Information Technology
Core stackAWS Glue · PySpark · Snowflake
StatusOpen to AWS data roles
02

Selected experience

01
Nov 2024 — PresentHouston, TX

AWS Data Engineer

High Radius Technology
AWS GlueRedshiftSnowflake
  • Supported end-to-end AWS Glue pipelines for invoices, payments, remittances, deductions, collections, and customer-account data.
  • Integrated ERP, CRM, bank, MySQL, Cassandra, API, and file data into curated Amazon Redshift and Snowflake layers.
  • Built PySpark, Spark SQL, Lambda, EMR, Databricks, Airflow/MWAA, and Step Functions workflows with retries, alerts, and recovery procedures.
  • Established CI/CD for data pipelines using CodePipeline, CodeBuild, Jenkins, Git, automated tests, and controlled release workflows.
02
Jan 2023 — Oct 2024Springfield, IL

AWS Data Engineer

State of Illinois
Public SectorEMRDatabricks
  • Implemented public-sector data solutions using Hadoop, Cloudera, Spark, AWS services, Amazon Redshift, and Snowflake.
  • Migrated on-premises Oracle ETL workloads to Amazon EMR, Databricks, Redshift, and Snowflake while preserving source-to-target controls.
  • Protected government datasets with AWS IAM and Secrets Manager while optimizing PySpark and Spark SQL jobs through partitioning, caching, and join tuning.
03
Oct 2020 — Dec 2022Athens, GA

Data Engineer

InnovaCare Health
Healthcare DataPySparkAWS Glue
  • Built AWS Glue pipelines for member, provider, eligibility, claims, encounter, and clinical-quality data from disparate healthcare systems.
  • Prepared source-to-target mappings, data-flow rules, and architecture documents for ingestion, transformation, security, lineage, and recovery.
  • Provisioned EMR and Databricks for secure batch and continuous processing, then built SQL components for healthcare staging, audit, transformation, and reporting.
04
Apr 2018 — Sep 2020McLean, VA

Big Data Developer

Giant Food
HadoopHiveMapReduce
  • Created Hadoop data solutions for retail POS, product, pricing, promotion, inventory, loyalty, pharmacy, and store operations data.
  • Installed and configured Hive, Pig, Sqoop, Flume, Oozie, HBase, Impala, Storm, Solr, and supporting Hadoop components.
  • Implemented Spark and Java MapReduce processing for XML, JSON, CSV, compressed files, logs, and retail feeds, with Cloudera-based operational support.
05
Jun 2016 — Mar 2018Chicago, IL

ETL Developer

Morningstar, Inc.
InformaticaOracleMarket Data
  • Translated investment-data requirements into source-to-target ETL specifications, technical designs, and delivery estimates.
  • Designed Informatica mappings, sessions, workflows, UNIX scripts, and parameter files for market, security, fund, pricing, and reference data.
  • Built reconciliation and UAT controls, and created SCD Type 1 and Type 2 dimensions and star schemas for investment analytics.
06
Jan 2014 — May 2016Baltimore, MD

ETL Developer

Transamerica
Data WarehouseInformaticaSSIS
  • Implemented enterprise data-warehouse solutions for insurance, retirement, policy, premium, claims, customer, and financial reporting data.
  • Built Informatica PowerCenter mappings and Mapplets to cleanse and load data into Operational Data Store and reporting structures.
  • Designed complex Informatica mappings and delivered SSIS, SSRS, VBA, and controlled flat-file components for insurance and retirement reporting.
12+
Years across cloud, big data, ETL, and warehouse delivery
6
Domains: finance, government, healthcare, retail, investment, insurance
AWS
Glue, S3, EMR, Redshift, Lambda, Step Functions, Kinesis, IAM
ETL
Batch, streaming, orchestration, data quality, CI/CD, production support
03

Engineering case studies

Finance SaaS · Cloud Data Platform

Receivables Data Lake & Finance Pipelines

Designed and supported AWS Glue, Lambda, EMR, Databricks, Airflow/MWAA, Step Functions, Redshift, and Snowflake workflows for invoices, payments, remittances, deductions, collections, customer accounts, ERP, CRM, bank, API, and secure-file sources.

AWS
Cloud-native pipelines
CI/CD
Controlled releases
DQ
Audit and reconciliation
AWS GlueLambdaEMR SnowflakeRedshiftKafka
Government · Data Modernization

Public-Sector Data Lake Modernization

Migrated on-premises Oracle ETL workloads and file-based feeds into Amazon S3, EMR, Databricks, Redshift, and Snowflake while preserving source-to-target controls, service identities, protected-data access, batch schedules, and reporting requirements.

Hybrid
Cloud and Hadoop
IAM
Protected access
Amazon S3EMRDatabricksHive
Healthcare · Claims Data Engineering

Claims, Eligibility & Provider Pipelines

Built AWS Glue, EMR, Databricks, DynamoDB, PySpark, Spark Streaming, and SQL workflows for member, provider, eligibility, claims, encounter, and clinical-quality data with security, lineage, recovery, and downstream interface documentation.

PHI
Protected workflows
Spark
Batch and streaming
AWS GluePySparkEMRDynamoDB
04

Core competencies

Cloud data architecture · ETL/ELT · Big data processing

Data platforms built for scale, controls, and production support

Pipelines
ETL/ELT, batch, streaming, orchestration
Platforms
AWS, Hadoop, Spark, Snowflake, Redshift
Warehousing
Data marts, ODS, star and snowflake schemas
Governance
Validation, lineage, security, source-to-target mapping
05

Technical toolkit

01 / Architecture

Cloud data platforms

AWS data lakes, analytical warehouses, secure services, and hybrid modernization.

02 / Pipelines

ETL and ELT delivery

Batch, event-driven, and streaming ingestion with source-to-target controls.

03 / Processing

Spark engineering

PySpark, Spark SQL, Scala, partitioning, caching, joins, and performance tuning.

04 / Warehousing

Data modeling

ODS, data marts, dimensional models, star schemas, and SCD patterns.

05 / Reliability

Data quality

Validation, reconciliation, lineage, audit controls, monitoring, and recovery.

06 / Delivery

Orchestration and CI/CD

Airflow, Step Functions, CodePipeline, CodeBuild, Jenkins, Git, and release controls.

Cloud/ 01
AWS GlueAmazon S3Redshift RDSAthenaDynamoDB EMRLambdaStep Functions KinesisCloudWatchIAM Secrets ManagerDatabricks
Programming/ 02
PythonPySparkScala Shell ScriptPerl Script SQLJavaPL/SQL
Big Data/ 03
HadoopHDFSMapReduce YARNSparkSpark SQL Spark StreamingSqoopHive OozieKafkaFlume ImpalaStormSolrPig
Warehouse, ETL & BI/ 04
InformaticaSSISSSAS SSRSSnowflakeOracle MySQLSQL ServerPostgreSQL MongoDBCassandraQuickSight TableauPower BIGitHub
06

Education & foundation

Education/ Degree
Master's in Information Technology
Catholic University of America
Aug 2012 — Dec 2013
Bachelor of Computer Science and Technology
Lovely Professional University
Aug 2008 — Apr 2012
Methods/ Delivery
Agile, Scrum, SDLC, CI/CD
Delivery methodology
Data governance, quality assurance, and production support
Operational controls
07

Delivery scope

Finance SaaS

Receivables & Payments

Invoices, remittances, deductions, collections, reconciliation, and customer-account data.
AWS Glue · Redshift · Snowflake · Kafka
Public Sector

Agency Data Platforms

Government datasets, protected access, Oracle migration, analytical warehouses, and reporting feeds.
S3 · EMR · Databricks · IAM
Healthcare

Claims & Eligibility

Member, provider, claims, encounter, clinical-quality, lineage, security, and recovery workflows.
AWS Glue · PySpark · EMR · DynamoDB
Retail

Store Operations Data

POS, product, pricing, promotion, inventory, loyalty, pharmacy, and transaction processing.
Hadoop · Hive · MapReduce · Spark
Investment Data

Market & Reference Data

Security, fund, portfolio, pricing, reference data, star schemas, SCD Type 1 and Type 2.
Informatica · Oracle · UNIX
Insurance

Policy & Claims Warehousing

Policy, premium, claims, customer, financial reporting, ODS loads, SSIS, and SSRS support.
Informatica · SQL Server · SSIS
Let's build reliable data platforms

Need cloud pipelines, warehouses,
or production data systems?

Open to AWS Data Engineer and enterprise data engineering opportunities in the United States.

File drop
Send a project brief, sample file, or dataset note directly to my inbox folder.
Private Console
The real DeepSeek API key stays in Cloudflare secrets. This browser only stores your private PIN.
PIN saved on this device
Ask DeepSeek anything.
Files upload through your Google Apps Script endpoint into the selected Drive folder. Visitors do not need to sign in with Google.
Not connected
Drop files here, or click to choose
Your DeepSeek API key belongs in Cloudflare secrets. This page only stores a private access PIN and calls the same-origin proxy endpoint.