Terraform nobody wants to touch
State drift, tangled modules, and a quiet fear that the next apply wakes the incident channel.
Owner-operated infrastructure studio
RuntimeForge is Dean Whitlock’s studio for AWS platforms, Terraform, inherited-environment recovery, and post-incident hardening. You brief one senior operator — and that operator reads the logs, makes the changes, and writes the handoff.
§ 01 — The problem
None of it is negligence. Systems accrete under deadlines until nobody can safely say what will happen when they change. RuntimeForge exists to make that legible again — and to fix it without a ceremonial rewrite.
State drift, tangled modules, and a quiet fear that the next apply wakes the incident channel.
IAM sprawl, long-lived keys, and permissions that grew faster than anyone’s ability to reason about them.
Production still leans on tribal memory and a few careful people staying awake at the right moment.
The fire is out, the postmortem is filed, and the actual remediation is still vague and unscheduled.
Every diagram and readout on this site is a labeled demonstration. Nothing here is real client telemetry, and no metrics, outcomes, or logos are invented.
§ 02 — Engagements
Sold like a studio, delivered like an operator. Scope is shaped around what is actually broken, risky, or expensive to leave vague — not around a fixed package.
Startups moving past improvisation
When production still depends on a few careful people and tribal memory, this pass turns it into infrastructure with repeatable deploys, environment separation, and clear operating boundaries.
AWS account structure, IAM baseline, Terraform module shape, CI/CD, secrets handling, DNS, and the first honest runbook.
Growth-stage teams carrying confusing environments
For the system that technically works but nobody trusts: drift, ambiguous access, hidden coupling, and a steady fear that the next edit wakes the incident channel.
Drift reconciliation, state and module untangling, access review, delivery hardening, and a prioritized remediation sequence.
Teams carrying fresh risk or an uncomfortable postmortem
Contain the immediate problem, review access and exposure, then implement hardening that actually lowers the chance of a repeat — without forcing a rewrite.
Exposure review, credential rotation, least-privilege enforcement, logging and detection gaps, and change-path safety.
§ 03 — Method
Most infrastructure damage happens during well-intentioned changes made without a clear picture. The work moves from investigation to safe implementation to a real handoff — in that order, every time.
Map the environment, find what is actually risky, and separate symptoms from the structural problem before touching anything.
Choose the safest path, make rollback visible, and keep urgency from quietly becoming permanent architecture.
The same person who did the diagnosis makes the changes. Detail does not get diluted through a handoff.
Runbooks, notes, and post-change context are part of the work — not optional cleanup that never happens.
§ 04 — Selected work
No client logos, no fabricated outcomes. The verifiable proof is a habit of reducing ambiguity until a system can be operated, inspected, and handed off with less guesswork.
A contained environment for traffic and patterns you do not want to generate in a real estate.
Mar 2026
A Rust networking build for when the router is technically up and operationally unhelpful.
Jan 2026
PowerShell utilities collapsing repetitive infrastructure work into a smaller set of commands.
Jun 2025
A TypeScript experiment where structure, sequencing, and visual systems get to be expressive.
Jan 2026
Custom Merlin firmware for edge hardware that still deserved maintenance and a better posture.
Oct 2023
§ 05 — Field notes
Incident patterns, infrastructure mistakes, and hardening work — the kind of detail that stays useful to buyers, operators, and whoever inherits the system next.
A persistence-by-alias pattern that hides in plain sight in the Google Workspace admin console, and the Admin SDK Apps Script that cleaned it up in under a minute.
Open noteA field guide to the environment-isolation failures that show up six months into a Terraform rollout, from shared secrets to DNS drift.
Open notePulumi, CDK, Ansible, plain shell — and the project shapes where Terraform is actually the wrong pick. A working taxonomy for picking the right IaC tool.
Open note§ 06 — Why direct
I’m Dean Whitlock. RuntimeForge is my studio. The work lives in AWS accounts, Terraform state, IAM sprawl, delivery pipelines, and the awkward edges that appear once infrastructure becomes consequential.
You are not buying project management wrapped around a specialist. You are buying direct technical judgment, implementation, and handoff from the same operator — which keeps infrastructure work moving when ambiguity and stakes climb.
Tooling matters, but not as identity. The useful thing is choosing the approach that fits the system, then leaving it clearer, safer, and easier to operate than before.
Working vocabulary
§ 07 — Contact
Scoped builds, cleanup, hardening passes, audits, and exploratory calls all fit here. Specific detail helps more than polished detail.
Message received
Expect a reply within one business day.
Form relay not configured in this environment
The form still validates locally. For live sends, set PUBLIC_FORMSPREE_ID, or use a direct channel.