Responsibilities
- Design, deploy, and maintain multi-cluster, multi-region Apache Kafka environments managed by the Strimzi Operator or similar on Kubernetes, as well as enterprise HiveMQ MQTT brokers.
- Drive the platform upgrade lifecycle (minor and major versions) for Kafka, Strimzi, and HiveMQ, coordinating platform patching cadences seamlessly to mitigate CVEs with zero application downtime.
- Develop and maintain custom Kafka Connectors (Source/Sink) for seamless data movement between legacy systems and the cloud.
- Formulate and implement scaling strategies, node affinity rules, topology spread constraints, and persistent volume management optimized for IOPS-intensive message workloads.
- Manage Kafka broker partition layouts, topic compaction configurations, segment retention behaviors, and cluster-wide resource allocation.
- Design high-availability and disaster recovery strategies, including cross-region replication and automated failover, to maintain platform resilience as usage scales.
- Enforce data governance standards through the management of Schema Registries (Avro, Protobuf, JSON Schema).
- Architect edge traffic routing to messaging clusters using advanced ingress layers, translating complex routing policies and managing edge capabilities like Envoy Gateway or custom proxies.
- Enforce data governance boundaries by implementing Kafka ACLs, SASL authentication structures, and integrating API Managers for authentication intercept handles (e.g., managing OAuth/OIDC validation paths)
- Implement robust end-to-end security frameworks, including mutual TLS (mTLS) validation, SSL verification, SASL/SCRAM and RBAC for the platforms.
- Build and maintain deep observability pipelines using Prometheus and Grafana, configuring custom dashboards to monitor critical broker KPIs.
- Formulate proactive alerting rules to catch cluster anomalies—such as KubeNodeNotReady, JVM memory degradation, disk capacity thresholds, or storage detached states—before they impact production workloads.
- Perform Root Cause Analysis (RCA) for platform outages and implement automated guardrails to prevent recurrence.
- Provide Tier 3 support for complex integration issues, such as consumer lag, rebalance loops, or network bottlenecks.
- Act as a technical subject matter expert (SME) for application engineering teams developing in Java / Spring Boot, advising on optimal Producer/Consumer settings (e.g., acks, idempotency, batch sizing, schema management).
- Ensure stability, performance, scalability, and cost efficiency
- Architect and maintain reusable GitHub Actions CI/CD pipelines to validate, lint, and test infrastructure manifests, Kafka client configurations, and custom plug-ins prior to deployment.
- Utilize automated scaling frameworks like KEDA (Kubernetes Event-driven Autoscaling) to dynamically scale application consumers based on real-time lag and throughput metrics.
- Treat infrastructure strictly as code by leveraging Helm for packaging, resource templating, and managing OCI artifacts.
- Provide clean abstractions, developer self-service tooling, and documented READMEs to empower development teams to provision topics, schemas, and credentials safely within architectural guardrails.
- Handle incident, problem, and change management at platform level.
- Work with Security, Architecture, and development teams at ICA.
Required Skills
Platform & Cloud
- 4–6 years of experience with Apache Kafka, Strimzi, Kafka Connect API, Kraft.
- 4–6 years of experience working with Java and Spring Boot Framework.
- Experience with PostgreSQL.
- Strong experience with Azure Kubernetes Service (AKS).
- Linux experience (including WSL).
- GitHub Actions and ArgoCD.
- Knowledge of Splunk, Fluent Bit and Helm.
DevOps & Automation
- CI/CD using GitHub Actions and/or Jenkins.
- Infrastructure as Code (Terraform or equivalent).
Observability
- Splunk, Grafana, Prometheus.
- Fluent Bit or OpenTelemetry.
Streaming & Messaging
- Apache Kafka at platform level (clusters, performance, security).
- HiveMQ broker.
Desired Skills
- Experience in APIM platforms.
- Experience in Azure, AWS or Google cloud.
- Cloud cost optimization experience.
- Experience of working at an Integration department in large enterprise environments.