Senior DevOps & Platform Engineer (Internal Developer Platform, DevSecOps, AIOps)
Location: Remote within EU
We are looking for a Senior / Lead Platform & DevOps Engineer with 8+ years of experience to design, build and scale an Internal Developer Platform, enterprise GitOps delivery and automated remediation across multi-cluster Kubernetes environments. The role sits in Integration Testing & Release DevOps and requires strong hands-on depth from day one.
Responsibilities
- Design and standardize a self-service Internal Developer Platform (e.g. Backstage or Port) with golden deployment paths and secure, ephemeral environments
- Architect multi-cluster, multi-region GitOps delivery with ArgoCD or Flux, including automated canary analysis, traffic shadowing and automated rollbacks
- Build closed-loop remediation and AIOps pipelines using OpenTelemetry, anomaly correlation and workflow/orchestration engines (Temporal, Kubernetes Operators in Go)
- Implement policy as code and software supply chain security with Kyverno or OPA, SAST/SCA gates, cosign and SBOMs
- Define SLO/SLI governance, error budgets and automated chaos testing for distributed services
- Establish standards for IaC, CI/CD and telemetry, and drive them through cross-team RFCs and architecture reviews
- Support inference and training infrastructure for AI workloads, including GPU autoscaling (KEDA)
Required experience
- 6+ years of enterprise-scale Kubernetes in production: multi-cluster networking, service meshes (Istio/Linkerd), ingress controllers, CRD design
- Strong Go (custom operators, platform controllers) and Python (automation, data and AI pipelines)
- Modular Terraform, OpenTofu or Pulumi, including state management, drift detection and policy testing
- GitOps and delivery tooling: ArgoCD, Flux, GitHub Actions
- Observability: OpenTelemetry collector topologies, tail-based sampling, Prometheus, Grafana (Mimir/Tempo), ClickHouse or Datadog
- Policy as code and supply chain security: Kyverno, OPA, SAST/SCA, cosign, SBOMs
- Deterministic, idempotent workflows with Temporal or event-driven architectures on Kafka
- Track record of authoring cross-team RFCs and defining operational SLOs/SLIs
- Cloud: AWS and/or GCP
- Incident tooling: PagerDuty, Incident.io, Slack API
Nice to have
- Integration of ITSM/CMDB platforms (iTop or similar) into automated incident workflows
- Experience integrating LLM tool-calling, agentic triage or RAG-based runbook retrieval into incident response
- AI inference infrastructure (Ray, vLLM, Triton, Kubeflow)
- Contributions to CNCF/open-source projects