Deutsche Telekom IT Solutions

Sovereign Engineering Platform SRE - T Cloud Public (REF5740Q)

Budapest, Szeged, Pécs, Debrecen, hu · Posted 5d ago

salary not listedDept: T Cloud Public
KubernetesHelmTerraformAnsibleArgoCDCrossplaneGitLabJenkinsvLLMOllamaQdrantOpen WebUIPrometheusGrafanaLokiOpenTelemetryELKOpenSearchVaultKeycloak

Job Description

Mission 
Design, build, and operate the secure infrastructure foundation used by Meridian engineering teams for AI-assisted software development, model experimentation, repository analysis, CI/CD execution, and controlled handover work in isolated or sovereignty-sensitive environments 

Role focus 
This infrastructure and operations role centers on Kubernetes-based engineering platforms, GitOps, private registries, internal model endpoints, observability, access control, and reliable operations for AI-enabled SDLC workloads. The candidate should enable engineering velocity while preserving security, auditability, and operational discipline 

Key responsibilities 

  • Build and operate Kubernetes environments that host AI engineering tools, internal model gateways, retrieval components, workflow services, CI/CD runners, and documentation services 

  • Implement GitOps and Infrastructure as Code patterns for reproducible provisioning, configuration, policy enforcement, platform upgrades, and disaster recovery readiness 

  • Manage private registries, package mirrors, secrets, identity integration, network segmentation, storage classes, backup routines, and controlled connectivity models 

  • Provide observability for engineering workloads, including metrics, logs, traces, GPU and CPU utilization, service health, cost signals, and operational runbooks 

  • Work with software, security, and architecture teams to ensure the platform supports AI-assisted SDLC workflows without creating uncontrolled data exposure or audit gaps 

Examples of market tools, models, and platform components expected 

  • Platform tooling such as Kubernetes, Helm, Terraform, Ansible, ArgoCD, Crossplane, GitLab runners, Jenkins agents, private registries, and internal package mirrors. 

  • AI platform components such as vLLM, Ollama, OpenAI-compatible gateways, Qdrant or similar vector stores, Open WebUI, Continue-compatible endpoints, and workflow services. 

  • Observability and operations stacks such as Prometheus, Grafana, Loki, OpenTelemetry, ELK/OpenSearch, Alertmanager, SRE runbooks, and incident management tooling. 

  • Security and governance components such as Vault, Keycloak, network policies, RBAC, admission controls, image scanning, SBOM tooling, and audit logging. 

  • Infrastructure awareness covering GPU-backed nodes, CPU-only fallback, storage performance, network isolation, proxy patterns, on-premise environments, and dedicated landing zones. 

Candidate profile 

  • 5+ years in SRE, platform engineering, DevOps, cloud infrastructure, or operations roles with strong Kubernetes and Linux expertise. 

  • Proven experience building and operating production-grade engineering platforms with GitOps, Infrastructure as Code, observability, and operational runbooks. 

  • Hands-on skills in Terraform, Ansible, Helm, Python or shell scripting, CI/CD runners, private registries, and secure configuration management. 

  • Good understanding of networking, storage, secrets, access control, monitoring, backup, disaster recovery, and operational hardening in high-security environments. 

  • Comfortable supporting AI-enabled engineering workloads in sovereignty-driven contexts where isolation, controlled data handling, reliability, and auditability are mandatory. 

Please note: remote working is only possible from within Hungary due to European taxation regulations.

* Please be informed that our remote working possibility is only available within Hungary due to European taxation regulation.