Allianz

Data Engineer for AI (f/m/d)

Frankfurt, Germany · Posted 4h ago

salary not listedpermanent
DatabricksDeltaUnity Catalog

Job Description

 

In the role of Data Engineer for AI (f/m/d), you will design, build and operate data products and reusable data preparation components on the Databricks-based Data and AI Platform at Allianz Global Investors.

 

Your focus is to enable platform customers to reliably source, ingest, transform, validate and serve high-quality, compliant data for AI and ML use cases (including analytics and GenAI), so teams can consume data in a self-service manner. You will provide technical guidance and L3 support to delivery teams, define and promote best practices for data pipelines and data quality, and ensure adherence to security, privacy and regulatory requirements following a compliance-by-design approach.

 

In addition, you will continuously evolve the platform’s data engineering capabilities by integrating new features (e.g., ingestion patterns, transformation frameworks, governance controls and monitoring) and by delivering standardized, reusable pipelines and templates that scale across use cases.

 

This position will be based in Frankfurt.

 

What you will do

 

  • Develop, maintain and enhance standardized pipeline patterns, templates and utilities (ingestion, transformation, validation, enrichment) to deliver AI/ML‑ready datasets
  • Build robust pipelines for structured and unstructured data with reproducible, deterministic outputs
  • Build and operate reliable batch and streaming pipelines on Databricks to ingest data from internal and external sources, curate datasets, and publish trusted data products for AI, ML and analytical consumption
  • Curated data products (Bronze/Silver/Gold): Establish controlled dataset evolution, enforce data quality and freshness, and curate gold tables explicitly designed for analytical, ML and AI use cases
  • Design, build and operate feature engineering pipelines and curate a ready‑to‑use feature store with proper versioning and lightweight documentation (ownership, purpose, inputs/outputs, SLAs)
  • Collaborate with ML and AI Engineers to build and operate pipelines for GenAI/RAG systems, including new data ingestion, Bronze/Silver updates, chunking, embedding refresh and index updates to maintain knowledge base relevance
  • Enable systematic experimentation (offline evaluations, A/B tests) and support model retraining using user feedback (binary and non‑binary) to continuously improve models
  • Collaborate with the Databricks Platform Architect to introduce and operationalize platform capabilities for ingestion, transformation and serving (Delta/Unity Catalog patterns, orchestration, monitoring), making them available securely and consistently
  • Facilitate provisioning and rollout of Databricks Lakeflow connectors to integrate new data sources and standardize ingestion across domains
  • Implement and enforce governance using Unity Catalog, including access control, lineage, retention, encryption and auditing, ensuring compliance with internal standards and regulations
  • Establish automated data quality checks and validations (e.g., DQX), define SLAs/SLOs for freshness and reliability, and implement monitoring and observability
  • Optimize Spark workloads, storage layouts and compute usage to ensure scalable, stable and cost‑efficient data pipelines
  • Partner with Data Scientists, ML Engineers, AI CoE, DevOps and SecOps to define data requirements, promote reusable patterns, and enable self‑service through documentation, coaching and reviews
  • Provide L3 support for pipeline and data product incidents, perform root‑cause analysis, and drive continuous improvements to platform stability and customer satisfaction
  • Integrate Databricks with (Azure) cloud services, enterprise systems, external providers and SaaS platforms, and expose curated data products securely to downstream consumers (feature stores, model training/inference, BI)

 

What you bring

 

Required:

  • Minimum 3 years of practical experience building and operating data pipelines and data products at scale on Databricks (batch and/or streaming), ideally in environments with strong governance requirements
  • Strong understanding of data engineering concepts for AI/ML (data modelling, data quality, data contracts, feature engineering collaboration, reproducibility) and how data characteristics impact model outcomes
  • Expert-level technology skills in Databricks, Spark, Python and SQL, plus solid engineering practices such as CI/CD and infrastructure-as-code (e.g., Terraform) on Azure
  • Diploma in computer science or a similar field
  • Very good understanding of data architecture and platform patterns (Lakehouse concepts, Delta, medallion approaches, data product thinking) and how to operationalize them in a governed enterprise context
  • Excellent analytical, planning, and organizational skills
  • Good oral and written communication skills in English
  • Ability to efficiently and effectively document and share knowledge and enable others

 

Preferred:

  • Working experience in the financial industry, preferably in asset management
  • International work experience
  • Certification “Databricks Certified Data Engineer Professional” (or equivalent)
  • Certification “Databricks Certified Data Engineer Associate” (or equivalent)
  • Optional: Certification “Databricks Certified Generative AI Engineer Associate” (helpful for GenAI-related data preparation patterns)

 

What we offer

 

  • We empower our employees by ensuring flexible work arrangements that maintain a balance between performance, productivity, career development and personal priorities (e.g., hybrid model/ flexible working hours)
  • Securing your future: Access to company pension/savings plans
  • Family support (relocation/ childcare facilities)
  • Company share purchasing plan
  • Mental health and wellbeing programs
  • Mobility solutions (Jobrad bike leasing, subvention Jobticket)
  • Career opportunities within the entire Allianz Group
  • Self-guided learning & development
  • Volunteering time
  • … and so much more!