Death By A Thousand Alarms
Taming alert fatigue, and the three levels of observability maturity.
ReadA technology leader with 15+ years across software development, developer tooling, reliability, and infrastructure.
By the numbers
About
I bring experience across multiple verticals — SaaS, fintech, healthcare, and hospitality — but I'm always open to jumping into something new. I'd love the opportunity to dive into HPC, developer tooling, B2B SaaS, and anything relating to SPACE.
Outside of work I'm an avid writer, sharing my thoughts on technology on my Medium blog. I mix, master, write, and record music, and I occasionally build guitars.
I reside in Herndon, VA with my amazing wife and three wacky kids.
I've built and led teams across the full breadth of engineering and operations — from software delivery to the network floor.
Areas of expertise
Growing engineers and building teams that ship — with <1% turnover.
Turning multi-year vision into roadmaps teams can actually execute.
Designing software and infrastructure that scales and stays maintainable.
Tooling and platforms that make the right thing the easy thing.
Leadership philosophy
State your intent and ask for input — don't ask me for permission to do your job. I'll give you the context to make good calls and the room to make them. Failure that comes from a well-reasoned decision is a learning opportunity; failure that comes from waiting to be told what to do is a management problem I need to fix.
I'll never share publicly what we haven't discussed privately. When something needs correcting, it happens in a one-on-one, close to the event, with a plan. Public praise, private correction — always, without exception.
Correction should come with a path forward, not just a verdict. If I'm pointing out something you did wrong, I'm also telling you what “right” looks like, why it matters, and what I'll do to help you get there. If I can't do that, I haven't thought about it hard enough yet.
You don't need permission to take your kid to the doctor, leave early for a school pickup, or block off time to think. If you're reachable and your team knows where you are, I trust you to manage your own time. After-hours messages wait until tomorrow unless they're marked otherwise. Time isn't the metric; output is.
I expect you to spend roughly a fifth of your time on things that make you better — side projects, self-directed learning, reading, writing, or exploring adjacent domains. This isn't a perk. It's how you stay sharp, how the team stays ahead, and how I ensure the people who work with me keep growing whether they stay on my team or eventually move on.
I own my team's failures the same way I celebrate their successes. If leadership needs to know something went wrong, I'm the one who tells them, and I'm the one whose name is on the fix. Your job is to do the work well; my job is to defend the conditions that let you do it.
Career highlights
Developed AI agents in Copilot Studio to ingest, triage, track, and disseminate real-time CVE feeds from the National Vulnerability Database. Agentic AI filtered, gauged CVE impact, and compiled multiple similar CVEs into comprehensive reporting packages.
Developed branching and delegation Claude skills as part of a larger SRE-focused incident triage pipeline.
Reduced public cloud expenditure by $1.3M annually through resource rightsizing, decommissioning stale resources, CSP and third-party recommendations, application re-architecture, and realistic data lifecycle policies.
Built and began executing a 5-year consolidation from seven cloud providers to two (AWS & Azure), migrating core AWS infrastructure to Terraform for IaC — targeting a net 5% cost reduction and far more manageable infrastructure.
With a strong focus on reliability, observability, and automation, increased product uptime by 3% annually — roughly 12 additional days of availability each year.
Increased team velocity by 27% by spearheading Agile-inspired processes in the SRE organization, lifting team morale and execution on strategic goals.
Reduced incident duration by 2% month-over-month through better application observability, production readiness, and post-incident Root Cause Analysis.
Cut build and deploy times by 30% and removed manual steps by designing and deploying a net-new CI/CD system focused on containerization, pipeline reusability, and end-to-end automated deploy and rollback.
Featured writing
Taming alert fatigue, and the three levels of observability maturity.
Read
The compensation, culture, and growth investments that make top talent stay.
Read
Real culture shows up in leadership's actions, not mission statements.
Read
Reflections on moving from hands-on technical work into people leadership.
Read
Rejecting the cult of constant progress in favor of starting fresh.
Read
A sudden diagnosis, and what facing mortality teaches about what matters.
ReadSkills & certifications
Work with me
Seeking technical leadership positions at the Director+ level (ideally VP or C-level).
Not sure what to do next or how to improve your organization? I can help.
Book me for your next all-hands, team offsite, conference, or special event.