cristian.coca/ staff platform engineer

$ terraform plan -target=career

Cristian Coca Hermosel

It's not about rebuilding.
It's about evolving.

Plan: 5 to add, 11 years to change, 0 to destroy.

I evolve systems while they run. Staff Platform Engineer working on the parts of a platform that can't stop: deployments, infrastructure as code at scale, and the AI agents that now help run both.

uptime
11years in infra
from zero
2×platforms founded
systems led
5in production, end to end
published
4articles on infra + AI
01 / resources

Systems in production

+ led  ·  ~ contributed  ·  0 to destroy

+ project.blue_green led live

Releases in the morning, and nobody notices.

Transparent blue/green deployment on EKS

beforeLarge releases ran in a maintenance window every few weeks. Changes piled up, each release got riskier, and the risk kept the window in place.

builtA second production cluster stays powered off until a release needs it. The pipeline boots it, applies only what changed in the release branch, and wires three ephemeral public APIs (operations, driver and passenger) to it, so the operations web and the mobile apps can test the new version with real logins before anyone switches. Promotion moves the traffic weights, the testing APIs fall back to a 503 stub, and the old cluster stays up for twelve hours in case we roll back.

proofDevelopers ship from the same workflow they always used. I ran the swap both ways in production and checked three independent sources: zero 5xx. The first real release went out at noon on a workday, and it is now the team's default way to ship.

release cycle · idle
Production traffic enters through Route 53 and reaches the active cluster. During a release, ephemeral testing APIs point at the offline cluster running the new version; after promotion they fall back to a 503 stub. route53 · apigw production users apigw · testing operation · driver · passenger 503 stub 100% 0% cluster-blue active · v1 · prod cluster-green offline · powered off
09:00:00state=idle · active-slot=blue · testing apis → 503 stub

+ project.terragrunt_migration led live

Splitting a Terraform monolith without stopping it.

Monolith to modular Terragrunt, strangler-fig style

beforeOne Terraform state covered every environment and account. Plans were slow, every apply felt dangerous, and nothing could be reused.

builtA library of reusable modules and a multi-account Terragrunt layout, migrated piece by piece. Each step extended what already existed, with no big-bang rewrite and no window where the platform stood still.

proofThe logic changed and the running infrastructure didn't. The same module library now underpins every environment, from networking to observability to the AI services.

strangler · migration
The Terraform monolith and new Terragrunt micro-states share one state storage. Each side reads the other's outputs while slices move from the monolith to their own states. STATE STORAGE · S3 reads tg outputs reads monolith outputs monolith.tfstate100% of resources tg-state-1network tg-state-2eks tg-state-3data terraform monolithlegacy root module · live terragruntmicro-states · 0 of 3 slices
readymonolith owns everything · terragrunt layout ready, empty

+ project.infra_architect led live

infra_architect: an Agentic AI IaC System

Multi-agent orchestrator for infrastructure tickets

beforeTurning an infrastructure ticket into changes across several repositories needed one senior engineer holding the whole system in their head.

builtAn orchestrator that reads the ticket, plans the work, and hands each part to a specialist agent: Terragrunt, modules, Lambda functions, APIs, and one that writes the summary for humans. Each agent has a narrow job and narrow context, which is what makes its output reliable enough to review.

proofA ticket goes in and a complete multi-repo plan comes out. The engineer reviews a plan instead of reconstructing the system from scratch.

agents · idle
A Jira ticket goes to infra_architect, an Opus orchestrator. It runs four Sonnet operators in parallel, a Haiku planner consolidates their output, and a Sonnet communicator opens pull requests and updates Jira. All models run on Bedrock. jira ticketinput infra_architectopus · orchestratoridle op_terragruntsonnetop_modulessonnetop_apisonnetop_serverlesssonnet infra_plannerhaiku · plan.md communicatorsonnet pull requestsgithub plan commentjira bedrock opus · sonnet · haiku · one model per job
readywaiting for a ticket

+ project.devsecops_remediation led live

DevSecOps Agentic AI System

Autonomous vulnerability remediation pipeline

beforeFour scanners (cloud posture, code, dependencies, IaC) produced four separate streams of alerts, and senior engineers sorted false positives by hand.

builtEvery finding lands in one source of truth. Before any model sees it, the pipeline attaches the real Terraform source and repository metadata. The model runs behind circuit breakers and token budgets, and its output forks into a pull request for safe fixes or a diagnostic report for anything that needs judgment.

proofA person reviews every change before it merges, and a remediation cycle now costs $0.35 instead of $75.

remediation · idle
Four scanners feed DefectDojo. Each finding is enriched, classified and given its Terraform source before an LLM sees it. The LLM runs behind a circuit breaker and either opens a pull request or writes a diagnostic report, and a person reviews both. SASTCSPMSCAIaC defectdojosingle source of truth DETERMINISTIC · NO LLM YET enrich finding filter + classify attach tf source llmbreaker · budget pull requestsafe fix diagnosticneeds judgment human review · always
ready4 scanners connected · 0 findings in queue

+ project.hub_and_spoke led live

Every account isolated, one hub in control.

Multi-account AWS network, hub-and-spoke

beforeNo dedicated network design and no one to own it. Environments had to be split cleanly before anything else could be built on top of them.

builtOne AWS account per environment plus a central hub. Address ranges are planned with IPAM so nothing overlaps, and each account follows the same three tiers: public load balancers, private services and data, and NAT. Engineers reach the spokes through a single VPN in the hub, signed in with SSO, and their group decides which ranges they can route to.

proofSpokes have no direct route to each other, so a mistake in dev can't reach prod. Critical data in prod is copied every week to a backup vault in the hub with its own encryption keys, and the restore was verified end to end.

network · scenarios
A central hub account connects to dev, pre and prod accounts and to engineers on the VPN. Spokes have no direct links to each other. hub vpn · sso · ipam backup vault devaccount · 3 tiers preaccount · 3 tiers prodaccount · 3 tiers engineervpn client · sso

Contributions

~ update in-place · team work where I owned a part

~ project.wefox_hybrid_network contributed

Offices, datacenter and AWS on one network.

Global hybrid network at wefox

wefox's global network had two cores. The first was a FortiGate HA cluster in the Colt datacenter in Barcelona, the hub for site-to-site VPNs from the offices in Barcelona, Berlin, Hanover, Zurich and Rome, and for the global VPN used by remote staff. The second was an AWS Transit Gateway in the shared services account, which replaced VPC peering: every VPC isolated from the others and reachable from the offices through the hub.

The cores were linked by Direct Connect, with an AWS IPsec VPN as automatic failover: BGP moved traffic over in under five seconds and back again when the line recovered. The AWS side was managed entirely in Terraform. I was one of the main engineers building and running it, and an escalation point for network incidents.

hybrid network · core links
Offices in Switzerland, Germany, Spain and Italy connect to a datacenter FortiGate. The FortiGate links to an AWS Transit Gateway over Direct Connect, with an IPsec VPN as backup. The Transit Gateway connects three AWS accounts. direct connect · bgp ipsec · failover office · CHZurich offices · DEBerlin · Hanover office · ESBarcelona office · ITRome colt · bcnfortigate HAcore 1 awstransit gw · svccore 2 isolatedvpc isolatedvpc isolatedvpc
primarydirect connect · datacenter fortigate ⇄ transit gateway
failoveripsec takes over in under 5s if direct connect drops, and hands back on recovery
resultevery office reaches every vpc through the hub; vpcs stay isolated from each other

~ project.wefox_identity contributed

One login for everything.

Centralized identity and single sign-on at wefox

A single Active Directory domain stretched across the datacenter and AWS: two domain controllers on premises and two in AWS Directory Service, replicating between sites so either side could keep authenticating on its own.

One identity opened everything: workstation sessions, email, internal applications, databases, the VPN and the office Wi-Fi. For the Wi-Fi I built and documented the 802.1X proof of concept: domain credentials checked over RADIUS, with each user dropped into their VLAN automatically. I was one of the main engineers on the identity platform.

identity · one domain, two sites
Two domain controllers in the datacenter and two in AWS Directory Service replicate with each other and form one domain. Workstations, email, internal apps, databases and the VPN all sign in against that domain. DATACENTER AWS · DIRECTORY SERVICE replication DC 1on premises DC 2on premises DC 3managed AD DC 4managed AD one domainsingle sign-on · one identity per person workstations email internal apps databases vpn wi-fi · 802.1x
02 / state history

Eleven years, first infra hire more than once

  1. 2023 → now
    Staff Platform Engineer & DevOps LeadFounded the DevOps function from zero: multi-account AWS, EKS, the IaC library, AIOps and LLMOps systems.
    • Blue/green in production: the first release went out at noon
    • Started writing publicly about platform engineering and LLMOps
    • Promoted to Staff Platform Engineer & DevOps Lead
    • Founder again: first DevOps hire, the platform from zero
    Celering
  2. 2024
    Freelance Project: Cloud EngineerAn AWS platform from zero: networking and peering, ECS on Spot, ALB, WAF, DNS and Aurora PostgreSQL.
    • My first time deploying infrastructure with AWS CDK in Java
  3. 2020 → 2023
    Cloud & Infrastructure EngineerObservability for the move to cloud native, then a core contributor to the global hybrid network on AWS Transit Gateway.
    • Built the backup and disaster recovery solution side by side with AWS Professional Services
    • My first time building core cloud infrastructure: SSO, networking and multi-account AWS
    • Moved in-house, from consultant to the core infrastructure team
    • Joined to build observability for the move to cloud native
    wefox
  4. 2018 → 2020
    Lead Infrastructure & Systems EngineerFirst systems engineer. Full migration from on-premise to AWS.
    • My first time as an infrastructure founder: from on-premise to AWS
    ACAI Performance
  5. 2017 → 2018
    Systems Engineer → Tech LeadHybrid cloud and virtualization; led the cloud support team.
    • Promoted to Tech Lead of the cloud support team
    • My first contact with the cloud: Azure
    Mediacloud
  6. 2015 → 2017
    Systems AdministratorDistributed infrastructure for a national retail network.
    • My first job in infrastructure
    Conforama
03 / writing

Published notes

on LinkedIn
04 / contact

Ready to apply?

Open to Staff and Principal platform roles, fully remote. Happy to go deeper on any of these systems, or to compare notes on yours.

loading…