Private Cloud Infrastructure Consulting
I design, migrate, and operate private cloud infrastructure, including self-managed Kubernetes, for engineering teams who are done overpaying hyperscalers, without sacrificing reliability.
Live session
What I Offer
Low commitment
A prioritized report on the security, cost, and reliability gaps in your current setup, plus a walkthrough call. No remediation pressure.
Audit in progress
Core build
A production-ready cluster, hardened, monitored, and GitOps-managed, fully documented and handed off to your team.
Release
Highest value
Move off AWS, GCP, or Azure onto private infrastructure you control, with a clear plan and validation at every step, so your infrastructure cost becomes predictable, never a monthly surprise.
Migration plan
How I Work
Review your current infrastructure, map the security, cost, and reliability gaps, and agree on what to fix first.
Implement the fixes and the target setup: hardened, documented, and validated at every step, with nothing hidden.
Keep it running with monitoring, upkeep, and on-call support, or hand it off cleanly to your team.
AI-Assisted Operations
The goal: offload the repetitive parts to AI, things like reviews, triage, and first drafts, so your team's job is to approve the change, not write it from scratch. These are examples; every workflow is built around how your team actually works.
When a PR opens against main, an automated review runs against your repository, checking for inconsistencies, optimization gaps, security issues, and anti-patterns before a human looks at it.
PR review pipeline
Trigger
PR opened against main
14 files changed
AI Agent
Automated review runs
Scans the diff for risk, style, and correctness
Findings
2 issues flagged
Unpinned base image, N+1 query in OrderService.list()
Output
Comment posted to PR #482
Awaiting human approval
A new error in your observability platform triggers an automatic investigation that traces the context and reproduces the trigger. Once the root cause is found, it opens a PR with the fix and a full explanation for the team.
Error investigation pipeline
Trigger
New error captured
Alert ERR-1180 raised by the observability platform
AI Agent
Investigation runs
Correlates with deploy a91f3c2, reproduces the trigger
Findings
Root cause found
Nil pointer in the payment webhook handler
Output
Fix PR opened (#491)
Includes root-cause notes for the team
When PagerDuty or Opsgenie fires, recent deployments, logs, and internal docs get pulled into a Slack thread automatically, including similar past incidents and the runbook to follow. After resolution, the incident history feeds a first-draft root-cause analysis.
On-call assist pipeline
Trigger
PagerDuty alert fires
Alert PD-3391
AI Agent
Assistant gathers context
Recent deploys, logs, internal docs, similar past incidents
Findings
Runbook matched
Relevant playbook attached to the thread
Output
Summary posted to #incidents
First-draft RCA follows after resolution
Infrastructure-as-code changes are checked against frameworks like SOC 2, HIPAA, and ISO 27001 as part of the PR, with inline comments pointing to the exact violation and a compliant configuration to use instead.
Compliance check pipeline
Trigger
IaC change opened in PR
PR #507
AI Agent
Compliance agent scans diff
Checked against SOC 2, HIPAA, and ISO 27001
Findings
Violation detected
Missing encryption-at-rest on module.storage
Output
Compliant config suggested inline
Awaiting human approval
Route, payload, and schema changes in a PR automatically update the OpenAPI spec, Postman collections, and internal wikis, so documentation stays accurate without anyone remembering to update it by hand.
Docs sync pipeline
Trigger
Schema change merged
Order.status enum
AI Agent
Docs-sync agent runs
Detects drift across specs and wikis
Findings
Updates generated
openapi.yaml, Postman collection
Output
Docs stay accurate
No manual edits required
Ongoing Support
Basic
Someone watching the lights: monitoring and upkeep, a few hours a month for when something breaks.
Standard
Backups and multi-node support layered on Basic, for infrastructure that's grown a few more moving parts.
Priority SRE
1-hour guaranteed response, 24/7 on-call, quarterly reviews, for teams where downtime means lost revenue.
Book a 30-minute call. We'll talk through what's going on and whether I'm the right fit.
Cookies
We use essential cookies to run the site, and optional analytics (PostHog and Google Analytics) to understand usage. Choose what you allow. Privacy policy
Essential cookies stay on. Turn analytics on only if you are comfortable with PostHog and Google Analytics.