Vectoryx.ai GmbH

AI Platform Engineer Senior/Mid

Remote · Posted 3d ago

salary not listedseniorcontractremote
PythonFastAPILangChainKubernetesDockerGitCI/CDArgoFluxJava

Job Description

We're building an AI-native claims operation. The models can already do remarkable work — what makes them reliable enough to hand real operational decisions to is the platform around them. That platform is what you'll own. We're an agile, senior team automating a complex, communication-heavy insurance operation, starting with residential property claims. Our agents run live cases today. Your job is to make the infrastructure they run on fast, observable, secure and boringly reliable — and to build it on solid, portable foundations rather than locking us into any one vendor.Your tasksThe Python services that orchestrate our agentic pipeline (eg. FastAPI, LangChain), and the platform they run onOur production runtime — deployment, scaling, resilience, and keeping services healthy under real loadRelease management — a proper branching and release strategy in Git, environment promotion, semantic versioning, and safe rollouts: canary/progressive delivery, feature flags, and clean rollback when something's wrongCI/CD pipelines and infrastructure-as-code — fast, safe, repeatable delivery from commit to productionThe async processing backbone — queued, scheduled and event-driven claim handlingThe data layer — choosing and running the right persistence and building it to scale, not inheriting whatever's easiestObservability and ops: tracing, logging, metrics, alertingSecurity and data protection — how we handle and isolate sensitive claim data in a GDPR-heavy domainYour profileHave 5+ years in backend/platform/DevOps engineering, and think in systems, not just featuresHave real, hands-on experience building and operating production services on KubernetesHave own a release process end to end — you know how to structure Git for a team, promote code across environments safely, and ship to production frequentlyAre strong in Python and comfortable owning cloud infrastructure end to end — cloud-agnostic by instinct, wary of lock-in, fluent in the primitives rather than any one provider's product catalogueLive in Docker, CI/CDHave solid database depth — you've diagnosed a slow query under load and fixed the root cause, and you can pick and run the right datastoreTreat reliability, observability and security as first-class concernsAre comfortable being the platform owner in a small team — high autonomy, high ownershipNice to haveExperience with LLM-serving infrastructure, queues, or agentic/async workloadsGitOps / progressive-delivery tooling (Argo, Flux, or similar)Java experience alongside PythonInsurance, fintech, or another regulated domain; German (helpful, not required — we work in English)