Software engineer with expertise in backend and cloud-native development, Kubernetes and LLM application development.
Aktualisiert am 10.07.2026
Profil
Freiberufler / Selbstständiger
Remote-Arbeit
Verfügbar ab: 01.08.2026
Verfügbar zu: 100%
davon vor Ort: 100%
Agentic-AI
Software
Cloud
LLMOps
AWS
RAG & Semantic Search
Kubernetes
Azure
SystemArchitektur
Cloud Engineer
DevOps Engineer
Docker
Golang
Software-Entwicklung
Observability
Python
Distributed Systems
English
Fluent
German
Proficient

Einsatzorte

Einsatzorte

Heidelberg (+500km) Zürich (+50km)
Deutschland, Schweiz

möglich

Projekte

Projekte

11 months
2025-06 - 2026-04

AI/ML Platform Engineering

AI Engineer LLM as Backend SystemArchitektur Back-End ...
AI Engineer
Built a scalable AI/ML platform for enterprise workloads, with support for RAG pipelines, model routing/orchestration, fine-tuning, and content-moderation workflows. Delivered cross-language SDKs, internal CLI tools, and fully automated CI/CD pipelines to streamline AI adoption for product teams while optimizing for cost efficiency, security, and operational reliability.


Responsibilities:

  • Designed and implemented AI?first CI/CD pipelines (Jenkins, GitHub Actions, GitLab CI, Argo Workflows/ArgoCD) that integrate model training, validation, and deployment into release automation.
  • Built fine-tuning, embedding, and moderation pipelines using Azure OpenAI and Azure Cognitive Services enhancing domain relevance, safety compliance, and multilingual support across enterprise applications. 
  • Implemented document processing pipelines integrating Azure Document Intelligence and OCR to extract, structure, and route enterprise content into downstream RAG and analytics workflows.
  • Automated deployment, scaling, and lifecycle management of GenAI workloads using Argo Workflows, ArgoCD, Jenkins, and GitOps-based configuration management. Achieved faster release cycles and significantly reduced configuration drift and deployment failures.
  • Integrated AI tooling into Git workflows and code quality processes (pre?merge LLM checks, automated code review bots, Copilot evaluation & safe?use playbooks) for dev teams. 
  • Authored operational runbooks, playbooks and hands?on workshop curricula (prompt engineering, LLMOps, secure Copilot usage, CI/CD integration) and coached engineering/SRE teams. 
  • Enforced GDPR-compliant data handling, EU AI Act governance controls, and IAM-based access policies across all model endpoints and platform components.
  • Designed and implemented multi-tenant model-routing services across AWS Bedrock, Azure OpenAI and Google AI with dynamic LLM selection based on latency, token cost, throughput, and task-specific performance.
  • Built FastAPI microservices as the backend foundation for SDK APIs and internal platform services, exposing REST endpoints and supporting WebSocket-based streaming for real-time inference responses.
  • Delivered cross-language SDK packages in Python, TypeScript and Go, enabling product teams to integrate LLM capabilities with minimal boilerplate.

LLM as Backend SystemArchitektur Back-End Python Prompt-Engineering Azure OpenAI RAG ArgoCD Jenkins GitOps Go Kserve Knative Kuberntes Inference Azure Devops Kubernetes KNative KServe Observability AWS Bedrock TypeScript Helm
SAP
9 months
2024-09 - 2025-05

Multi-Agent Autonomous AI Platform

Senior GenAI Engineer Azure AI Foundry MCP LangGraph ...
Senior GenAI Engineer
Designed and built a production-grade multi-agent platform enabling autonomous execution of complex, multi-step business processes. The system used a layered architecture: a tool exposure layer built on MCP servers (implemented as async Python FastAPI microservices with REST and SSE APIs), an orchestration layer managing agent graphs and inter-agent communication, and an operations layer covering state persistence, observability, guardrails, and human-in-the-loop controls. Primary use cases included autonomous document analysis pipelines, process-level decision support, observability, guardrails, and human-in-the-loop controls. Primary use cases included autonomous document analysis pipelines, semantic-search-powered retrieval, process-level decision support, and cross-system task execution with structured audit trails.


Responsibilities: 

  • Designed the MCP server architecture: implemented multiple domain-specific MCP servers in async Python (FastAPI) exposing internal APIs, SQL databases, file systems, and external services as agent-callable tools over REST and SSE endpoints. 
  • Built an autonomous agentic orchestration layer with LangGraph, enabling LLM agents to plan, tool-select, and execute multi-step workflows via stateful graphs, retries, and composable subgraphs for resilient long-running processes. 
  • Implemented a multi-agent supervisor pattern coordinating specialized sub-agents (retrieval agent, analysis agent, action agent) with structured handoff messages and shared scratchpad state.
  • Integrated vector databases (pgvector, Chroma) for semantic search and knowledge retrieval within the retrieval agent, enabling high-relevance document lookup across enterprise content stores. 
  • Integrated state persistence for agent sessions, approval workflows, audit logs, and checkpoint history, ensuring full traceability and resumability. 
  • Integrated agents with multiple LLM providers via a unified model router, abstracting provider differences and enabling per-task model selection.
  • Built human-in-the-loop controls: interrupt nodes triggering async approval requests via webhook; resumption logic restoring full agent state after human confirmation or correction. 
  • Applied GDPR-compliant data handling and IAM-based access control across all agent tool interfaces and external integrations. 
  • Set up agent observability via open-source tools for trace capture, LLM evaluation, and explainability of agent decision paths. 
  • Managed infrastructure deployment of MCP servers and agent services via Docker and GitHub Actions; defined resource limits and horizontal scaling policies for agent execution workers on Kubernetes.

Azure AI Foundry MCP LangGraph A2A protocol Agentic RAG FastAPI async Python Vector Databases (pgvector Chroma) SQL REST / SSE / Webhooks LLM Evaluation Docker Kubernetes GDPR / IAM
Enterprise SaaS
9 months
2024-01 - 2024-09

Agentic AI Platform & RAG Excellence Framework

Senior GenAI Engineer Python OpenAI LangChain ...
Senior GenAI Engineer

Took ownership of enterprise-wide Azure AI Foundry platform operations for a regulated environment, covering full lifecycle management of LLM and embedding model deployments, workspace governance, and internal platform adoption across multiple product teams. Governance frameworks were aligned with EU AI Act, ISO 27001, and DORA requirements. Extended the platform with Semantic Kernel-based, vector database integration for RAG-based knowledge retrieval for internal process orchestration.


Responsibilities:

  • Administered and maintained Azure AI Foundry workspaces across dev, test, and production environments, enforcing consistent workspace structure, compute configurations, and model catalog standards.
  • Managed end-to-end LLM and embedding model deployment lifecycle within Azure AI Foundry, including version control, staged rollouts, deprecation workflows, and quota allocation across business units.
  • Configured and operated Portkey as the AI gateway layer for model routing, request caching, semantic guardrails, and unified observability across AI endpoints.
  • Integrated Azure AI Gateway for token-based rate limiting, quota enforcement, semantic caching, and load balancing across model deployments, with AI Content Safety integration for automated prompt moderation and compliance.
  • Evaluated and integrated Semantic Kernel as an agent skill orchestration layer, enabling reusable, composable AI capabilities across internal product teams.
  • Integrated vector databases for RAG-based knowledge retrieval, enabling semantic search across enterprise document repositories and internal knowledge bases.
  • Implemented governance and compliance controls aligned with EU AI Act, ISO 27001, and DORA requirements, including audit logging, model lineage tracking, and explainability tooling for regulated AI deployments.
  • Configured IAM-based access policies and enforced secure data flows and GDPR-compliant data handling across all Azure AI Foundry workspace components.
  • Designed and implemented reusable AI agents and tool interfaces through MCPs, exposing internal services via APIs and MCP-compatible patterns.
  • Built agent execution workflows, enabling LLM-driven systems to trigger and orchestrate real-world actions through structured service calls.
  • Designed and implemented enterprise system integrations, connecting AI platforms with internal business services and external APIs to enable real-world data access and actions.

Python OpenAI LangChain Kubernetes RAG Chatbot Go (Golang) Firebase Firestore LlamaIndex Semantic Kernel Vector Databases Redis MCP AWS FastAPI async asynchronous programming Azure AI Foundry Azure AI Gateway Portkey Hybrid Retrieval Databricks Azure RBAC / IAM GitOps EU AI Act / ISO 27001 / DORA
SAP
Remote
1 year
2023-02 - 2024-01

Observability Engineering (AI/ML)

Observability Engineer (AI/ML) Prometheus Grafana Kubernetes ...
Observability Engineer (AI/ML)

Designed and implemented a fully instrumented, cloud-native observability and telemetry framework for hosted, fine-tuned, and proxied AI/ML models in enterprise-grade production environments. Delivered end-to-end visibility into AI/ML training pipelines, inference workloads, and model serving infrastructure.


Responsibilities:

  • Architected end-to-end observability pipelines for ML APIs and model-serving runtimes using OpenTelemetry SDKs/collectors, Prometheus exporters, and Kubernetes operators, instrumenting the full lifecycle of model training, inference, and system-level resource utilization.
  • Instrumented model endpoints, batch/stream training jobs, and inference gateways to capture high-resolution metrics such as tail latency, throughput (RPS/QPS), token-per-second performance, GPU memory fragmentation, multi-node utilization, error budgets, and anomaly detection signals.
  • Profiled and monitored inference optimization for vLLM and TensorRT-LLM deployments, tracking CUDA kernel performance, NCCL communication overhead, and memory bandwidth utilization to identify and resolve latency regressions.
  • Monitored distributed multi-GPU training runs across nodes connected via InfiniBand, capturing per-GPU utilization, gradient synchronization bottlenecks, and HPC cluster health signals for large-scale model training workloads.
  • Implemented automated alerting and SLO/SLA monitoring using Prometheus Alertmanager and custom anomaly-detection pipelines to identify inference latency regressions, GPU/CPU saturation events, memory leaks, container restarts, or failed model-training runs.
  • Collaborated with MLOps, SRE, and platform engineering teams to integrate telemetry into CI/CD pipelines, automate environment drift detection, and enable data-driven scaling policies for training and inference clusters.

Prometheus Grafana Kubernetes Helm Python OpenTelemetry Promitor Dynatrace Go Loki Jaeger Tempo KEDA HelmArgoCD MLflow ArgoCD vLLM TensorRT-LLM CUDA / NCCL Multi-GPU / HPC
SAP
11 months
2022-05 - 2023-03

Development of a distributed orchestration system

Senior software engineer Go(lang) WebSocket OpenSearch ...
Senior software engineer
Built a custom event-driven container-orchestration and control-plane platform inspired by Kubernetes using asynchronous processing, message queues, and streaming patterns to enable automated cloud resource provisioning for hundreds of products across thousands of tenants. Infrastructure provisioning was automated using AWS CDK for infrastructure-as-code, with workloads deployed across AWS Lambda, ECS, and S3. Client SDKs were delivered in both Go and TypeScript with code-generation tooling.


Responsibilities:

  • Designed and implemented the core control-plane architecture, including an API Server with RBAC, Controller Manager, Scheduler-like reconciliation loops, Namespace isolation, and CRD-style resource definitions.
  • Developed client SDKs and code-generation tools in Go and TypeScript to streamline custom controller development for internal engineering teams.
  • Provisioned and managed AWS infrastructure using AWS CDK (infrastructure-as-code), deploying workloads across AWS Lambda (serverless event handlers), ECS (containerized services), and S3 (object storage backends).
  • Implemented SQL/ORM-based resource state tracking and tenant isolation across the control plane data layer, supporting consistent reads and optimistic concurrency for thousands of concurrent tenants.
  • Integrated WebSocket for real-time event streaming between control plane and worker nodes and utilized Redis for distributed caching and message brokering.
  • Employed LocalStack for local AWS service emulation in CI/CD pipelines.
  • Integrated Prometheus, Alertmanager, and Grafana for end-to-end metrics instrumentation, distributed tracing, and proactive anomaly detection.
  • Implemented Kubernetes-native patterns: Deployments, StatefulSets, DaemonSets, custom operators.

Go(lang) WebSocket OpenSearch LocalStack Redis Prometheus Grafana AWS S3 Kubernetes RBAC Distributed Systems Webhook Go (Golang) TypeScript Microservices SQL AWS (Lambda ECS S3 CDK) Kubernetes
SAP
Walldorf
1 year 8 months
2020-09 - 2022-04

Enablement of SAP Analytics Cloud?s SaaS offering

Senior cloud engineer Cloud Foundry Node.js Java ...
Senior cloud engineer
Built core monetization and platform services for SAP Analytics Cloud, designing and delivering billing, metering, and service-broker capabilities that enabled scalable, pay-as-you-go SaaS consumption across hundreds of tenants.


Responsibilities:

  • esigned and implemented a Cloud Foundry service broker (Node.js), enabling automated provisioning, binding, and lifecycle management of platform services.
  • Architected and developed billing and metering microservices to capture, aggregate, and process resource usage data, forming the foundation for usage-based pricing and chargeback models.
  • Designed event-driven and API-based communication between microservices, ensuring reliable data flow and consistency across billing pipelines.
  • Built scalable data processing and storage layers (PostgreSQL, Redis) to handle high-volume usage metrics and support concurrent tenant workloads.
  • Implemented end-to-end observability using Prometheus and Grafana, including metrics instrumentation, alerting, and dashboards to ensure system reliability and performance.
  • Designed and exposed internal and external APIs to support seamless integration with platform services and customer-facing systems.
  • Ensured horizontal scalability and resilience of services, supporting hundreds of concurrent tenants with consistent performance
Cloud Foundry Node.js Java Prometheus API gateway Redis Postgres Grafana
Walldorf
1 year 5 months
2019-04 - 2020-08

Developed cloud infra. for the SAP HANA-as-a-Service

DevOps engineer HashiCorp (Terraform/ Vault/ Consul) Ansible AWS (VPC/ EC2/ S3/ Glacier/ Cloud Watch/ API Gateway) ...
DevOps engineer
Engineered a scalable, multi-region cloud infrastructure platform for SAP HANA-as-a-Service on AWS, enabling fully automated provisioning, upgrades, and lifecycle management of enterprise database systems. Designed the system to meet high availability, security, and compliance requirements while significantly accelerating deployment speed.


Responsibilities:

  • Architected and implemented infrastructure-as-code using Terraform and configuration management with Ansible to automate SAP HANA installation, upgrades, and system lifecycle operations across environments.
  • Designed and developed an event-driven automation layer, including a lightweight agent integrated with Consul for service discovery and change propagation, triggering dynamic infrastructure workflows.
  • Built and exposed APIs to handle customer HANA system provisioning and lifecycle operations, enabling self-service and programmatic access.
  • Engineered secure, production-grade AWS network architecture (VPC, subnets, routing, IAM), ensuring isolation, compliance, and high availability across regions.
  • Implemented automated backup and disaster recovery strategies leveraging S3 and Glacier, ensuring data durability and rapid restoration.
  • Reduced end-to-end deployment time by 40% through automation and system optimization while maintaining strict enterprise security and compliance standards.
HashiCorp (Terraform/ Vault/ Consul) Ansible AWS (VPC/ EC2/ S3/ Glacier/ Cloud Watch/ API Gateway) Cloud Foundry Python Go (Golang) Bash HashiCorp (Terraform Vault Consul) AWS (VPC EC2 S3 Glacier Cloud Watch API Gateway)
SAP
10 months
2018-07 - 2019-04

Development of an elastic caching microservice

Software Engineer Go(lang) Redis MongoDB ...
Software Engineer
Developed an elastic caching microservice in Go to accelerate analytical query performance for a multi-tenant analytics platform, implementing context-aware caching with user permissions, roles, and cube dimension metadata.


Responsibilities:

  • Built cloud-native microservice following 12-factor app principles with stateless design, externalized configuration, and graceful shutdown
  • Implemented multi-layer caching strategy using Redis (L1 in-memory cache with TTL/LRU eviction) and MongoDB (L2 persistent cache for complex query metadata)
  • Designed cache key generation algorithm incorporating RBAC permissions, tenant isolation, and OLAP cube context (dimensions, measures, filters)
  • Developed cache invalidation strategies with pub/sub patterns for real-time data updates
  • Integrated Prometheus with custom metrics (cache hit/miss ratios, query latency percentiles, eviction rates)
  • Built Grafana dashboards for real-time performance monitoring and capacity planning.
Go(lang) Redis MongoDB Kubernetes Prometheus 12-factor app
SAP

Aus- und Weiterbildung

Aus- und Weiterbildung

2014 - 2017
Distributed Software Systems
TU Darmstadt (Germany)
Degree: Master of Science

Position

Position

Software and DevOps engineer with focus on cloud-native development and LLM application development.

Kompetenzen

Kompetenzen

Top-Skills

Agentic-AI Software Cloud LLMOps AWS RAG & Semantic Search Kubernetes Azure SystemArchitektur Cloud Engineer DevOps Engineer Docker Golang Software-Entwicklung Observability Python Distributed Systems

Schwerpunkte

AI/ML Platform Engineering & LLMOps
Experte
Cloud-Native Observability & SRE
Experte
Distributed Systems & Platform Engineering
Fortgeschritten

AI/ML Platform Engineering & LLMOps

Deep expertise in building production-grade GenAI platforms and agentic AI systems, with comprehensive experience in LLM deployment, fine-tuning, RAG pipelines, and model orchestration. Specialized in architecting multi-tenant AI infrastructure that balances performance, cost optimization, and enterprise security requirements.


Cloud-Native Observability and SRE

Expert in designing end-to-end observability solutions for distributed systems and AI/ML workloads using OpenTelemetry, Prometheus, and Grafana. Proven ability to instrument complex environments from token-level metrics to infrastructure telemetry, enabling proactive incident management, anomaly detection, and data-driven optimization of high-throughput systems.


Distributed Systems and Platform Engineering

Strong foundation in building scalable, cloud-native platforms with expertise in Kubernetes ecosystem, control-plane architecture, and microservices orchestration. Skilled in implementing GitOps workflows, CI/CD automation, and infrastructure-as-code practices to deliver reliable, self-service platforms for enterprise-scale deployments.

Aufgabenbereiche

System Architecture
Experte
Software Engineering
Experte
DevOps
Fortgeschritten
  • Architecture and implementation of enterprise GenAI platforms supporting RAG, fine-tuning, model routing, and content moderation workflows
  • Design and deployment of observability frameworks for AI/ML systems, including distributed tracing, metrics pipelines, and SLO/SLA monitoring
  • Development of autonomous agent systems with tool integration, memory modules, and safety/governance controls
  • Building cloud-native microservices and APIs for multi-tenant SaaS offerings with focus on scalability and reliability
  • Infrastructure automation using GitOps, CI/CD pipelines, and infrastructure-as-code across AWS and Azure environments
  • Implementation of control-plane architectures for container orchestration and resource provisioning at scale
  • Establishment of LLMOps practices including experiment tracking, model versioning, drift detection, and compliance enforcement
  • Performance optimization through caching strategies, autoscaling policies, and resource utilization monitoring
  • Cross-functional collaboration with ML Ops, SRE, and platform engineering teams to accelerate AI adoption
  • Security implementation including RBAC, multi-tenant isolation, and content moderation pipelines

Produkte / Standards / Erfahrungen / Methoden

DevOps
Experte
Software
Experte
AWS
Fortgeschritten
OpenAI
Experte
Kubernetes
Experte
Observability
Fortgeschritten
GenAI
Experte
Development
Experte
PromptFlow
Azure API Management
Microsoft Power Platform
Azure Storage
Azure Queue
GitLab CI
Design Patterns
gRPC
Profile
As a freelance software engineer, I deliver tailored, scalable software solutions for enterprise systems. My focus spans software architecture and development, DevOps and LLM-powered AI applications, with strong expertise in building reliable, observable and cost-efficient distributed systems on AWS and Azure.

AI/ML & GenAI
OpenAI, Azure OpenAI, LangChain, LlamaIndex, Semantic Kernel, RAG (Retrieval-Augmented Generation), Prompt Engineering, LLM Fine-tuning, Model Inference, Vector Databases, MLflow, Kserve, Knative, MCP (Model Context Protocol), Chatbot Development, Agentic AI Systems


Cloud Platforms & Services

AWS (VPC, EC2, S3, Glacier, CloudWatch, API Gateway), Azure DevOps, Azure Cognitive Services, Cloud Foundry, Multi-cloud Architecture


Container Orchestration & Infrastructure

Kubernetes, Helm, ArgoCD, Argo Workflows, Docker, StatefulSets, DaemonSets, Custom Operators, Control Plane Architecture


Observability & Monitoring

OpenTelemetry, Prometheus, Grafana, Dynatrace, Promitor, Loki, Jaeger, Tempo, Alertmanager, Distributed Tracing, Metrics Engineering, SLO/SLA Monitoring


Programming Languages

Python, Go (Golang), Node.js, Java, Bash


Data Storage & Caching

Redis, MongoDB, PostgreSQL, Firebase Firestore, Vector Databases, OpenSearch, AWS S3


DevOps & Automation

GitOps, Jenkins, Terraform, Ansible, HashiCorp (Vault, Consul, Terraform), LocalStack, CI/CD Pipelines, GitHub Actions, Infrastructure-as-Code (IaC)


Networking & Communication

REST APIs, WebSocket, API Gateway, RBAC, Service Mesh


Development Practices & Patterns

Microservices Architecture, 12-Factor App Principles, SRE Practices, LLMOps, MLOps, Multi-tenant Design, Distributed Systems, Event-Driven Architecture, KEDA (Kubernetes Event-Driven Autoscaling)


Security & Compliance

RBAC (Role-Based Access Control), Content Moderation, Prompt Injection Defense, Multi-tenant Isolation, Policy Enforcement, Adversarial Testing


Data & ML Tools

DVC (Data Version Control), Weights & Biases, Model Registries, Experiment Tracking, Dataset Versioning

Betriebssysteme

Linux

Programmiersprachen

Go (Golang)
Python
Java
Node.js
Postgres
MongoDB
Firebase

Datenbanken

PostgresSQL
MongoDB
Redis
Firestore
Elasticsearch

Branchen

Branchen

  • Enterprise Software

  • SaaS

Einsatzorte

Einsatzorte

Heidelberg (+500km) Zürich (+50km)
Deutschland, Schweiz

möglich

Projekte

Projekte

11 months
2025-06 - 2026-04

AI/ML Platform Engineering

AI Engineer LLM as Backend SystemArchitektur Back-End ...
AI Engineer
Built a scalable AI/ML platform for enterprise workloads, with support for RAG pipelines, model routing/orchestration, fine-tuning, and content-moderation workflows. Delivered cross-language SDKs, internal CLI tools, and fully automated CI/CD pipelines to streamline AI adoption for product teams while optimizing for cost efficiency, security, and operational reliability.


Responsibilities:

  • Designed and implemented AI?first CI/CD pipelines (Jenkins, GitHub Actions, GitLab CI, Argo Workflows/ArgoCD) that integrate model training, validation, and deployment into release automation.
  • Built fine-tuning, embedding, and moderation pipelines using Azure OpenAI and Azure Cognitive Services enhancing domain relevance, safety compliance, and multilingual support across enterprise applications. 
  • Implemented document processing pipelines integrating Azure Document Intelligence and OCR to extract, structure, and route enterprise content into downstream RAG and analytics workflows.
  • Automated deployment, scaling, and lifecycle management of GenAI workloads using Argo Workflows, ArgoCD, Jenkins, and GitOps-based configuration management. Achieved faster release cycles and significantly reduced configuration drift and deployment failures.
  • Integrated AI tooling into Git workflows and code quality processes (pre?merge LLM checks, automated code review bots, Copilot evaluation & safe?use playbooks) for dev teams. 
  • Authored operational runbooks, playbooks and hands?on workshop curricula (prompt engineering, LLMOps, secure Copilot usage, CI/CD integration) and coached engineering/SRE teams. 
  • Enforced GDPR-compliant data handling, EU AI Act governance controls, and IAM-based access policies across all model endpoints and platform components.
  • Designed and implemented multi-tenant model-routing services across AWS Bedrock, Azure OpenAI and Google AI with dynamic LLM selection based on latency, token cost, throughput, and task-specific performance.
  • Built FastAPI microservices as the backend foundation for SDK APIs and internal platform services, exposing REST endpoints and supporting WebSocket-based streaming for real-time inference responses.
  • Delivered cross-language SDK packages in Python, TypeScript and Go, enabling product teams to integrate LLM capabilities with minimal boilerplate.

LLM as Backend SystemArchitektur Back-End Python Prompt-Engineering Azure OpenAI RAG ArgoCD Jenkins GitOps Go Kserve Knative Kuberntes Inference Azure Devops Kubernetes KNative KServe Observability AWS Bedrock TypeScript Helm
SAP
9 months
2024-09 - 2025-05

Multi-Agent Autonomous AI Platform

Senior GenAI Engineer Azure AI Foundry MCP LangGraph ...
Senior GenAI Engineer
Designed and built a production-grade multi-agent platform enabling autonomous execution of complex, multi-step business processes. The system used a layered architecture: a tool exposure layer built on MCP servers (implemented as async Python FastAPI microservices with REST and SSE APIs), an orchestration layer managing agent graphs and inter-agent communication, and an operations layer covering state persistence, observability, guardrails, and human-in-the-loop controls. Primary use cases included autonomous document analysis pipelines, process-level decision support, observability, guardrails, and human-in-the-loop controls. Primary use cases included autonomous document analysis pipelines, semantic-search-powered retrieval, process-level decision support, and cross-system task execution with structured audit trails.


Responsibilities: 

  • Designed the MCP server architecture: implemented multiple domain-specific MCP servers in async Python (FastAPI) exposing internal APIs, SQL databases, file systems, and external services as agent-callable tools over REST and SSE endpoints. 
  • Built an autonomous agentic orchestration layer with LangGraph, enabling LLM agents to plan, tool-select, and execute multi-step workflows via stateful graphs, retries, and composable subgraphs for resilient long-running processes. 
  • Implemented a multi-agent supervisor pattern coordinating specialized sub-agents (retrieval agent, analysis agent, action agent) with structured handoff messages and shared scratchpad state.
  • Integrated vector databases (pgvector, Chroma) for semantic search and knowledge retrieval within the retrieval agent, enabling high-relevance document lookup across enterprise content stores. 
  • Integrated state persistence for agent sessions, approval workflows, audit logs, and checkpoint history, ensuring full traceability and resumability. 
  • Integrated agents with multiple LLM providers via a unified model router, abstracting provider differences and enabling per-task model selection.
  • Built human-in-the-loop controls: interrupt nodes triggering async approval requests via webhook; resumption logic restoring full agent state after human confirmation or correction. 
  • Applied GDPR-compliant data handling and IAM-based access control across all agent tool interfaces and external integrations. 
  • Set up agent observability via open-source tools for trace capture, LLM evaluation, and explainability of agent decision paths. 
  • Managed infrastructure deployment of MCP servers and agent services via Docker and GitHub Actions; defined resource limits and horizontal scaling policies for agent execution workers on Kubernetes.

Azure AI Foundry MCP LangGraph A2A protocol Agentic RAG FastAPI async Python Vector Databases (pgvector Chroma) SQL REST / SSE / Webhooks LLM Evaluation Docker Kubernetes GDPR / IAM
Enterprise SaaS
9 months
2024-01 - 2024-09

Agentic AI Platform & RAG Excellence Framework

Senior GenAI Engineer Python OpenAI LangChain ...
Senior GenAI Engineer

Took ownership of enterprise-wide Azure AI Foundry platform operations for a regulated environment, covering full lifecycle management of LLM and embedding model deployments, workspace governance, and internal platform adoption across multiple product teams. Governance frameworks were aligned with EU AI Act, ISO 27001, and DORA requirements. Extended the platform with Semantic Kernel-based, vector database integration for RAG-based knowledge retrieval for internal process orchestration.


Responsibilities:

  • Administered and maintained Azure AI Foundry workspaces across dev, test, and production environments, enforcing consistent workspace structure, compute configurations, and model catalog standards.
  • Managed end-to-end LLM and embedding model deployment lifecycle within Azure AI Foundry, including version control, staged rollouts, deprecation workflows, and quota allocation across business units.
  • Configured and operated Portkey as the AI gateway layer for model routing, request caching, semantic guardrails, and unified observability across AI endpoints.
  • Integrated Azure AI Gateway for token-based rate limiting, quota enforcement, semantic caching, and load balancing across model deployments, with AI Content Safety integration for automated prompt moderation and compliance.
  • Evaluated and integrated Semantic Kernel as an agent skill orchestration layer, enabling reusable, composable AI capabilities across internal product teams.
  • Integrated vector databases for RAG-based knowledge retrieval, enabling semantic search across enterprise document repositories and internal knowledge bases.
  • Implemented governance and compliance controls aligned with EU AI Act, ISO 27001, and DORA requirements, including audit logging, model lineage tracking, and explainability tooling for regulated AI deployments.
  • Configured IAM-based access policies and enforced secure data flows and GDPR-compliant data handling across all Azure AI Foundry workspace components.
  • Designed and implemented reusable AI agents and tool interfaces through MCPs, exposing internal services via APIs and MCP-compatible patterns.
  • Built agent execution workflows, enabling LLM-driven systems to trigger and orchestrate real-world actions through structured service calls.
  • Designed and implemented enterprise system integrations, connecting AI platforms with internal business services and external APIs to enable real-world data access and actions.

Python OpenAI LangChain Kubernetes RAG Chatbot Go (Golang) Firebase Firestore LlamaIndex Semantic Kernel Vector Databases Redis MCP AWS FastAPI async asynchronous programming Azure AI Foundry Azure AI Gateway Portkey Hybrid Retrieval Databricks Azure RBAC / IAM GitOps EU AI Act / ISO 27001 / DORA
SAP
Remote
1 year
2023-02 - 2024-01

Observability Engineering (AI/ML)

Observability Engineer (AI/ML) Prometheus Grafana Kubernetes ...
Observability Engineer (AI/ML)

Designed and implemented a fully instrumented, cloud-native observability and telemetry framework for hosted, fine-tuned, and proxied AI/ML models in enterprise-grade production environments. Delivered end-to-end visibility into AI/ML training pipelines, inference workloads, and model serving infrastructure.


Responsibilities:

  • Architected end-to-end observability pipelines for ML APIs and model-serving runtimes using OpenTelemetry SDKs/collectors, Prometheus exporters, and Kubernetes operators, instrumenting the full lifecycle of model training, inference, and system-level resource utilization.
  • Instrumented model endpoints, batch/stream training jobs, and inference gateways to capture high-resolution metrics such as tail latency, throughput (RPS/QPS), token-per-second performance, GPU memory fragmentation, multi-node utilization, error budgets, and anomaly detection signals.
  • Profiled and monitored inference optimization for vLLM and TensorRT-LLM deployments, tracking CUDA kernel performance, NCCL communication overhead, and memory bandwidth utilization to identify and resolve latency regressions.
  • Monitored distributed multi-GPU training runs across nodes connected via InfiniBand, capturing per-GPU utilization, gradient synchronization bottlenecks, and HPC cluster health signals for large-scale model training workloads.
  • Implemented automated alerting and SLO/SLA monitoring using Prometheus Alertmanager and custom anomaly-detection pipelines to identify inference latency regressions, GPU/CPU saturation events, memory leaks, container restarts, or failed model-training runs.
  • Collaborated with MLOps, SRE, and platform engineering teams to integrate telemetry into CI/CD pipelines, automate environment drift detection, and enable data-driven scaling policies for training and inference clusters.

Prometheus Grafana Kubernetes Helm Python OpenTelemetry Promitor Dynatrace Go Loki Jaeger Tempo KEDA HelmArgoCD MLflow ArgoCD vLLM TensorRT-LLM CUDA / NCCL Multi-GPU / HPC
SAP
11 months
2022-05 - 2023-03

Development of a distributed orchestration system

Senior software engineer Go(lang) WebSocket OpenSearch ...
Senior software engineer
Built a custom event-driven container-orchestration and control-plane platform inspired by Kubernetes using asynchronous processing, message queues, and streaming patterns to enable automated cloud resource provisioning for hundreds of products across thousands of tenants. Infrastructure provisioning was automated using AWS CDK for infrastructure-as-code, with workloads deployed across AWS Lambda, ECS, and S3. Client SDKs were delivered in both Go and TypeScript with code-generation tooling.


Responsibilities:

  • Designed and implemented the core control-plane architecture, including an API Server with RBAC, Controller Manager, Scheduler-like reconciliation loops, Namespace isolation, and CRD-style resource definitions.
  • Developed client SDKs and code-generation tools in Go and TypeScript to streamline custom controller development for internal engineering teams.
  • Provisioned and managed AWS infrastructure using AWS CDK (infrastructure-as-code), deploying workloads across AWS Lambda (serverless event handlers), ECS (containerized services), and S3 (object storage backends).
  • Implemented SQL/ORM-based resource state tracking and tenant isolation across the control plane data layer, supporting consistent reads and optimistic concurrency for thousands of concurrent tenants.
  • Integrated WebSocket for real-time event streaming between control plane and worker nodes and utilized Redis for distributed caching and message brokering.
  • Employed LocalStack for local AWS service emulation in CI/CD pipelines.
  • Integrated Prometheus, Alertmanager, and Grafana for end-to-end metrics instrumentation, distributed tracing, and proactive anomaly detection.
  • Implemented Kubernetes-native patterns: Deployments, StatefulSets, DaemonSets, custom operators.

Go(lang) WebSocket OpenSearch LocalStack Redis Prometheus Grafana AWS S3 Kubernetes RBAC Distributed Systems Webhook Go (Golang) TypeScript Microservices SQL AWS (Lambda ECS S3 CDK) Kubernetes
SAP
Walldorf
1 year 8 months
2020-09 - 2022-04

Enablement of SAP Analytics Cloud?s SaaS offering

Senior cloud engineer Cloud Foundry Node.js Java ...
Senior cloud engineer
Built core monetization and platform services for SAP Analytics Cloud, designing and delivering billing, metering, and service-broker capabilities that enabled scalable, pay-as-you-go SaaS consumption across hundreds of tenants.


Responsibilities:

  • esigned and implemented a Cloud Foundry service broker (Node.js), enabling automated provisioning, binding, and lifecycle management of platform services.
  • Architected and developed billing and metering microservices to capture, aggregate, and process resource usage data, forming the foundation for usage-based pricing and chargeback models.
  • Designed event-driven and API-based communication between microservices, ensuring reliable data flow and consistency across billing pipelines.
  • Built scalable data processing and storage layers (PostgreSQL, Redis) to handle high-volume usage metrics and support concurrent tenant workloads.
  • Implemented end-to-end observability using Prometheus and Grafana, including metrics instrumentation, alerting, and dashboards to ensure system reliability and performance.
  • Designed and exposed internal and external APIs to support seamless integration with platform services and customer-facing systems.
  • Ensured horizontal scalability and resilience of services, supporting hundreds of concurrent tenants with consistent performance
Cloud Foundry Node.js Java Prometheus API gateway Redis Postgres Grafana
Walldorf
1 year 5 months
2019-04 - 2020-08

Developed cloud infra. for the SAP HANA-as-a-Service

DevOps engineer HashiCorp (Terraform/ Vault/ Consul) Ansible AWS (VPC/ EC2/ S3/ Glacier/ Cloud Watch/ API Gateway) ...
DevOps engineer
Engineered a scalable, multi-region cloud infrastructure platform for SAP HANA-as-a-Service on AWS, enabling fully automated provisioning, upgrades, and lifecycle management of enterprise database systems. Designed the system to meet high availability, security, and compliance requirements while significantly accelerating deployment speed.


Responsibilities:

  • Architected and implemented infrastructure-as-code using Terraform and configuration management with Ansible to automate SAP HANA installation, upgrades, and system lifecycle operations across environments.
  • Designed and developed an event-driven automation layer, including a lightweight agent integrated with Consul for service discovery and change propagation, triggering dynamic infrastructure workflows.
  • Built and exposed APIs to handle customer HANA system provisioning and lifecycle operations, enabling self-service and programmatic access.
  • Engineered secure, production-grade AWS network architecture (VPC, subnets, routing, IAM), ensuring isolation, compliance, and high availability across regions.
  • Implemented automated backup and disaster recovery strategies leveraging S3 and Glacier, ensuring data durability and rapid restoration.
  • Reduced end-to-end deployment time by 40% through automation and system optimization while maintaining strict enterprise security and compliance standards.
HashiCorp (Terraform/ Vault/ Consul) Ansible AWS (VPC/ EC2/ S3/ Glacier/ Cloud Watch/ API Gateway) Cloud Foundry Python Go (Golang) Bash HashiCorp (Terraform Vault Consul) AWS (VPC EC2 S3 Glacier Cloud Watch API Gateway)
SAP
10 months
2018-07 - 2019-04

Development of an elastic caching microservice

Software Engineer Go(lang) Redis MongoDB ...
Software Engineer
Developed an elastic caching microservice in Go to accelerate analytical query performance for a multi-tenant analytics platform, implementing context-aware caching with user permissions, roles, and cube dimension metadata.


Responsibilities:

  • Built cloud-native microservice following 12-factor app principles with stateless design, externalized configuration, and graceful shutdown
  • Implemented multi-layer caching strategy using Redis (L1 in-memory cache with TTL/LRU eviction) and MongoDB (L2 persistent cache for complex query metadata)
  • Designed cache key generation algorithm incorporating RBAC permissions, tenant isolation, and OLAP cube context (dimensions, measures, filters)
  • Developed cache invalidation strategies with pub/sub patterns for real-time data updates
  • Integrated Prometheus with custom metrics (cache hit/miss ratios, query latency percentiles, eviction rates)
  • Built Grafana dashboards for real-time performance monitoring and capacity planning.
Go(lang) Redis MongoDB Kubernetes Prometheus 12-factor app
SAP

Aus- und Weiterbildung

Aus- und Weiterbildung

2014 - 2017
Distributed Software Systems
TU Darmstadt (Germany)
Degree: Master of Science

Position

Position

Software and DevOps engineer with focus on cloud-native development and LLM application development.

Kompetenzen

Kompetenzen

Top-Skills

Agentic-AI Software Cloud LLMOps AWS RAG & Semantic Search Kubernetes Azure SystemArchitektur Cloud Engineer DevOps Engineer Docker Golang Software-Entwicklung Observability Python Distributed Systems

Schwerpunkte

AI/ML Platform Engineering & LLMOps
Experte
Cloud-Native Observability & SRE
Experte
Distributed Systems & Platform Engineering
Fortgeschritten

AI/ML Platform Engineering & LLMOps

Deep expertise in building production-grade GenAI platforms and agentic AI systems, with comprehensive experience in LLM deployment, fine-tuning, RAG pipelines, and model orchestration. Specialized in architecting multi-tenant AI infrastructure that balances performance, cost optimization, and enterprise security requirements.


Cloud-Native Observability and SRE

Expert in designing end-to-end observability solutions for distributed systems and AI/ML workloads using OpenTelemetry, Prometheus, and Grafana. Proven ability to instrument complex environments from token-level metrics to infrastructure telemetry, enabling proactive incident management, anomaly detection, and data-driven optimization of high-throughput systems.


Distributed Systems and Platform Engineering

Strong foundation in building scalable, cloud-native platforms with expertise in Kubernetes ecosystem, control-plane architecture, and microservices orchestration. Skilled in implementing GitOps workflows, CI/CD automation, and infrastructure-as-code practices to deliver reliable, self-service platforms for enterprise-scale deployments.

Aufgabenbereiche

System Architecture
Experte
Software Engineering
Experte
DevOps
Fortgeschritten
  • Architecture and implementation of enterprise GenAI platforms supporting RAG, fine-tuning, model routing, and content moderation workflows
  • Design and deployment of observability frameworks for AI/ML systems, including distributed tracing, metrics pipelines, and SLO/SLA monitoring
  • Development of autonomous agent systems with tool integration, memory modules, and safety/governance controls
  • Building cloud-native microservices and APIs for multi-tenant SaaS offerings with focus on scalability and reliability
  • Infrastructure automation using GitOps, CI/CD pipelines, and infrastructure-as-code across AWS and Azure environments
  • Implementation of control-plane architectures for container orchestration and resource provisioning at scale
  • Establishment of LLMOps practices including experiment tracking, model versioning, drift detection, and compliance enforcement
  • Performance optimization through caching strategies, autoscaling policies, and resource utilization monitoring
  • Cross-functional collaboration with ML Ops, SRE, and platform engineering teams to accelerate AI adoption
  • Security implementation including RBAC, multi-tenant isolation, and content moderation pipelines

Produkte / Standards / Erfahrungen / Methoden

DevOps
Experte
Software
Experte
AWS
Fortgeschritten
OpenAI
Experte
Kubernetes
Experte
Observability
Fortgeschritten
GenAI
Experte
Development
Experte
PromptFlow
Azure API Management
Microsoft Power Platform
Azure Storage
Azure Queue
GitLab CI
Design Patterns
gRPC
Profile
As a freelance software engineer, I deliver tailored, scalable software solutions for enterprise systems. My focus spans software architecture and development, DevOps and LLM-powered AI applications, with strong expertise in building reliable, observable and cost-efficient distributed systems on AWS and Azure.

AI/ML & GenAI
OpenAI, Azure OpenAI, LangChain, LlamaIndex, Semantic Kernel, RAG (Retrieval-Augmented Generation), Prompt Engineering, LLM Fine-tuning, Model Inference, Vector Databases, MLflow, Kserve, Knative, MCP (Model Context Protocol), Chatbot Development, Agentic AI Systems


Cloud Platforms & Services

AWS (VPC, EC2, S3, Glacier, CloudWatch, API Gateway), Azure DevOps, Azure Cognitive Services, Cloud Foundry, Multi-cloud Architecture


Container Orchestration & Infrastructure

Kubernetes, Helm, ArgoCD, Argo Workflows, Docker, StatefulSets, DaemonSets, Custom Operators, Control Plane Architecture


Observability & Monitoring

OpenTelemetry, Prometheus, Grafana, Dynatrace, Promitor, Loki, Jaeger, Tempo, Alertmanager, Distributed Tracing, Metrics Engineering, SLO/SLA Monitoring


Programming Languages

Python, Go (Golang), Node.js, Java, Bash


Data Storage & Caching

Redis, MongoDB, PostgreSQL, Firebase Firestore, Vector Databases, OpenSearch, AWS S3


DevOps & Automation

GitOps, Jenkins, Terraform, Ansible, HashiCorp (Vault, Consul, Terraform), LocalStack, CI/CD Pipelines, GitHub Actions, Infrastructure-as-Code (IaC)


Networking & Communication

REST APIs, WebSocket, API Gateway, RBAC, Service Mesh


Development Practices & Patterns

Microservices Architecture, 12-Factor App Principles, SRE Practices, LLMOps, MLOps, Multi-tenant Design, Distributed Systems, Event-Driven Architecture, KEDA (Kubernetes Event-Driven Autoscaling)


Security & Compliance

RBAC (Role-Based Access Control), Content Moderation, Prompt Injection Defense, Multi-tenant Isolation, Policy Enforcement, Adversarial Testing


Data & ML Tools

DVC (Data Version Control), Weights & Biases, Model Registries, Experiment Tracking, Dataset Versioning

Betriebssysteme

Linux

Programmiersprachen

Go (Golang)
Python
Java
Node.js
Postgres
MongoDB
Firebase

Datenbanken

PostgresSQL
MongoDB
Redis
Firestore
Elasticsearch

Branchen

Branchen

  • Enterprise Software

  • SaaS

Vertrauen Sie auf Randstad

Im Bereich Freelancing
Im Bereich Arbeitnehmerüberlassung / Personalvermittlung

Fragen?

Rufen Sie uns an +49 89 500316-300 oder schreiben Sie uns:

Das Freelancer-Portal

Direktester geht's nicht! Ganz einfach Freelancer finden und direkt Kontakt aufnehmen.