JOB DESCRIPTION
TERRAFORM / GCP DEVOPS ENGINEER (Cloud Infrastructure)
POSITION OVERVIEW
We are seeking an experienced Terraform/GCP DevOps Engineer to join our managed engineering team delivering infrastructure and cloud operations
You will be responsible for:
- Designing, building, and maintaining enterprise-grade infrastructure on Google Cloud Platform (GCP)
- Managing Infrastructure-as-Code (Terraform) across production environments
- Building and optimizing CI/CD pipelines for continuous deployment
- Ensuring platform reliability, security, and cost efficiency
- Leading disaster recovery, backup strategies, and business continuity
- Mentoring other engineers on infrastructure and cloud best practices
KEY RESPONSIBILITIES
Infrastructure Design & Architecture
- Design and architect scalable, resilient cloud infrastructure on GCP for production workloads
- Define infrastructure patterns, standards, and best practices for the platform
- Evaluate cloud services and tooling (Cloud Run, Cloud Functions, Firestore, BigQuery, Pub/Sub, etc.)
- Plan capacity and auto-scaling strategies to meet SLA requirements (99%+ availability)
- Design disaster recovery (DR) and business continuity (BC) architectures (RTO: 1 day, RPO: 24 hours)
- Document architecture decisions (ADRs) and maintain up-to-date system design documentation
- Conduct architecture reviews with engineering teams to ensure alignment with IKEA standards
Infrastructure-as-Code & Terraform
- Build and maintain Terraform modules for all GCP resources (compute, networking, storage, databases, security)
- Manage Terraform state securely with remote state backend (GCS + Terraform Cloud)
- Implement infrastructure versioning, testing, and code review workflows
- Automate infrastructure provisioning, scaling, and updates
- Ensure infrastructure code is modular, reusable, and well-documented
- Maintain 100% Infrastructure-as-Code coverage (zero manual changes in production)
- Plan and execute infrastructure updates with zero-downtime deployments
- Implement drift detection and auto-remediation for infrastructure state
CI/CD Pipeline & Deployment Automation
- Design and maintain GitHub Actions CI/CD pipelines for infrastructure and application deployments
- Implement automated testing for infrastructure code (Terraform validation, policy checks, security scanning)
- Build deployment automation and orchestration (Blue-green deployments, canary releases)
- Manage secrets and credentials securely (Google Secret Manager, encrypted Terraform variables)
- Implement automated rollback procedures and chaos engineering tests
- Monitor deployment quality metrics (change failure rate, mean time to recovery)
- Work with Cloudflare for edge deployment and traffic management
Cloud Operations & Reliability
- Manage GCP project structure, IAM policies, and security controls
- Implement monitoring, observability, and alerting across all infrastructure components
- Set up and optimize cloud resource monitoring (Cloud Monitoring, Cloud Logging, Trace)
- Manage cloud budgets, cost optimization, and resource efficiency
- Execute routine operational tasks: backups, maintenance windows, security patching
- Participate in on-call rotation for infrastructure incidents (P1/P2)
- Troubleshoot infrastructure issues and optimize performance
- Maintain detailed runbooks and disaster recovery procedures
Security, Compliance & Cost Management
- Implement infrastructure security best practices (network segmentation, encryption, IAM least privilege)
- Manage firewall rules, VPC configuration, Cloud Armor, and DDoS protection
- Ensure compliance with security standards, data residency, and privacy requirements
- Conduct infrastructure security audits and vulnerability assessments
- Manage data backup and encryption strategies
- Implement cost optimization strategies (reserved instances, committed use discounts, resource optimization)
- Monitor and reduce cloud spend while maintaining performance
- Manage cloud provider relationships and licensing
Mentoring & Knowledge Transfer
- Mentor junior/mid-level engineers on infrastructure, cloud platforms, and DevOps best practices
- Lead knowledge transfer sessions on Terraform, GCP, and CI/CD best practices
- Contribute to team documentation and internal knowledge base
- Participate in code reviews for infrastructure and deployment changes
- Lead design reviews and architecture discussions
REQUIRED QUALIFICATIONS
Experience
- 5+ years of professional DevOps/Cloud Infrastructure engineering experience
- 3+ years of hands-on experience with Terraform in production environments (preferably managing 50+ resources)
- 3+ years of production experience on Google Cloud Platform (GCP) or equivalent (AWS/Azure) with deep expertise in 2+ major GCP services
- 2+ years managing CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, or similar)
- Proven experience building and maintaining disaster recovery (DR) and business continuity (BC) strategies
- Experience operating customer-facing, high-availability services with 99%+ uptime targets
- Track record of incident response and on-call operations in production environments
Technical Skills - GCP & Cloud Platforms
Core GCP Services
- Compute: Deep expertise with Cloud Run (containerized workloads), Cloud Functions (serverless), Compute Engine
- Databases: Firestore (NoSQL), BigQuery (data warehouse), Cloud SQL (relational databases)
- Messaging & Streaming: Pub/Sub, Dataflow
- Storage: Cloud Storage (GCS), Firestore Backup & Point-in-Time Recovery (PITR)
- Networking: VPC, Cloud Load Balancer, Cloud Armor, VPN, Cloud Interconnect
- Security: Cloud IAM, Secret Manager, Cloud KMS, Security Command Center
- Monitoring & Observability: Cloud Monitoring, Cloud Logging, Cloud Trace, Profiler
- Deployment & Orchestration: Cloud Deploy, GKE (optional but valuable)
Infrastructure-as-Code & Terraform
- Advanced Terraform proficiency: modules, variables, outputs, state management, remote backends
- Terraform best practices: version control, testing, CI/CD integration
- Terraform Cloud / Enterprise (optional, valuable for remote state management)
- Terraform testing frameworks: Terratest, checkov, terraform validate, tflint
- Policy-as-Code: OPA/Conftest or similar for infrastructure policy enforcement
- Hands-on experience managing 50+ cloud resources with Terraform
- Git workflow and code review processes for infrastructure code
CI/CD & Deployment Automation
- GitHub Actions expertise (workflow design, custom actions, secrets management)
- GitHub Advanced Security: code scanning, dependency scanning, secret scanning
- Container technologies: Docker, container registries (Artifact Registry, Container Registry)
- Blue-green and canary deployment strategies
- Infrastructure-as-Code testing and validation in pipelines
- Automated rollback and disaster recovery procedures
- Monitoring deployment quality metrics
Observability & Monitoring
- OpenTelemetry instrumentation and integration
- Google Cloud Monitoring and Cloud Logging
- Sentry for error tracking and alerting
- Dashboard design and visualization (Grafana, Cloud Monitoring dashboards)
- Alert design, incident triage, and escalation procedures
- Metrics collection, tracing, and debugging distributed systems
Networking & Security
- VPC design, subnetting, routing, NAT/PAT
- Cloud Load Balancer and traffic management
- Cloud Armor (DDoS protection, WAF)
- VPN and encrypted communication
- IAM policy design and least-privilege access
- Secret management and credential rotation
- Network security auditing and compliance
Disaster Recovery & Backup
- Backup strategy design (RTO/RPO planning)
- Firestore backup and point-in-time recovery
- GCS versioning and retention policies
- Infrastructure disaster recovery testing
- Runbook development and documentation
PREFERRED QUALIFICATIONS
- Experience with Cloudflare (Workers, edge computing, WAF/DDoS)
- Kubernetes/GKE experience (container orchestration, Helm, operators)
- Experience with data pipelines (Dataflow, Cloud Composer/Airflow)
- Cost optimization expertise (RI analysis, committed use discounts, workload-based sizing)
- Experience with multi-region or multi-cloud deployments
- Background in platform engineering or building internal developer platforms (IDP)
- Exposure to governance and compliance (SOC 2, ISO 27001, GDPR)
- Experience with chaos engineering and resilience testing (Gremlin, Chaos Toolkit)
- Knowledge of API Gateway, Apigee or similar
- Experience with service mesh (Istio, Consul) - optional
- Terraform Cloud or Terraform Enterprise administration
- Linux/Unix system administration background