- I am a German Citizen with Indian roots.
- I have my own company in Germany
? Design and operate Delta Live Tables (DLT) pipelines on Azure Databricks for large-scale marketing automation, ingesting customer interaction and campaign event data into a medallion-architecture Delta Lake with declarative data quality expectations.
? Build rate-limited, fault-tolerant API delivery services that push audience segments and campaign triggers to external marketing and CRM systems ? implementing throttling, retry/backoff, and idempotency to respect downstream API quotas.
? Deploy all pipelines and jobs via Databricks Asset Bundles with CI/CD on Azure DevOps, enabling versioned, reviewable promotion across dev, test, and production workspaces.
? Develop ML models for campaign targeting and customer propensity scoring on Databricks with MLflow, feeding scored audiences directly into automated campaign journeys.
? Provision and govern the platform with Terraform (Databricks workspaces, ADLS Gen2, networking, secrets) and Unity Catalog for access control, lineage, and audit logging.
? Architected an end-to-end data platform on GCP with BigQuery, Dataflow, and real-time Kafka streaming handling 1M+ events/hour; deployed auto-scaling GKE clusters, reducing infrastructure costs by 40%.
? Rebuilt legacy ingestion on Databricks with Delta Lake and PySpark, achieving 2× faster batch execution and 99.9% pipeline uptime through robust error handling and CI/CD monitoring.
? Ingested real-time machine utilisation data from IoT sensors via Kafka to dynamically calculate credit risk premiums for industrial asset-backed loans ? a novel data product combining streaming architecture with risk modelling.
? Built an MLOps pipeline on Vertex AI with automated CI/CD for model training and deployment; fine-tuned LLAMA2 with LoRA for bank-specific document analysis processing 10K+ PDFs/day.
? Implemented an automated data quality framework with Git/Jenkins pipelines and deployed an Elasticsearch + Kibana stack for real-time fraud signal monitoring and log aggregation.
? Architected a Hive-based data warehouse with automated ETL pipelines processing 30M+ records per batch cycle; developed an interval-clustering anomaly detection algorithm for a ?15M+ recycling fraud detection platform (featured in The Guardian).
? Built cloud-native data pipelines on AWS using Lambda, Step Functions, and CodePipeline CI/CD; deployed an Elasticsearch search platform with automated NLP indexing, reducing resolution time by 60%.
? Designed real-time analytics dashboards in Tableau and created a Python-based compliance screening tool integrating the OpenCorporates API for sanctions exposure checks.
- I am a German Citizen with Indian roots.
- I have my own company in Germany
? Design and operate Delta Live Tables (DLT) pipelines on Azure Databricks for large-scale marketing automation, ingesting customer interaction and campaign event data into a medallion-architecture Delta Lake with declarative data quality expectations.
? Build rate-limited, fault-tolerant API delivery services that push audience segments and campaign triggers to external marketing and CRM systems ? implementing throttling, retry/backoff, and idempotency to respect downstream API quotas.
? Deploy all pipelines and jobs via Databricks Asset Bundles with CI/CD on Azure DevOps, enabling versioned, reviewable promotion across dev, test, and production workspaces.
? Develop ML models for campaign targeting and customer propensity scoring on Databricks with MLflow, feeding scored audiences directly into automated campaign journeys.
? Provision and govern the platform with Terraform (Databricks workspaces, ADLS Gen2, networking, secrets) and Unity Catalog for access control, lineage, and audit logging.
? Architected an end-to-end data platform on GCP with BigQuery, Dataflow, and real-time Kafka streaming handling 1M+ events/hour; deployed auto-scaling GKE clusters, reducing infrastructure costs by 40%.
? Rebuilt legacy ingestion on Databricks with Delta Lake and PySpark, achieving 2× faster batch execution and 99.9% pipeline uptime through robust error handling and CI/CD monitoring.
? Ingested real-time machine utilisation data from IoT sensors via Kafka to dynamically calculate credit risk premiums for industrial asset-backed loans ? a novel data product combining streaming architecture with risk modelling.
? Built an MLOps pipeline on Vertex AI with automated CI/CD for model training and deployment; fine-tuned LLAMA2 with LoRA for bank-specific document analysis processing 10K+ PDFs/day.
? Implemented an automated data quality framework with Git/Jenkins pipelines and deployed an Elasticsearch + Kibana stack for real-time fraud signal monitoring and log aggregation.
? Architected a Hive-based data warehouse with automated ETL pipelines processing 30M+ records per batch cycle; developed an interval-clustering anomaly detection algorithm for a ?15M+ recycling fraud detection platform (featured in The Guardian).
? Built cloud-native data pipelines on AWS using Lambda, Step Functions, and CodePipeline CI/CD; deployed an Elasticsearch search platform with automated NLP indexing, reducing resolution time by 60%.
? Designed real-time analytics dashboards in Tableau and created a Python-based compliance screening tool integrating the OpenCorporates API for sanctions exposure checks.