Experience
Senior cloud platform and site reliability engineer with 20+ years in AWS and Azure infrastructure, security, and automation. Datadog subject-matter expert building AI-driven automation for cloud security and observability in a regulated (HIPAA/PCI) environment. Outside the day job I run Three Moons Network, an independent practice that brings production-grade AI automation to small and medium businesses.
Download PDF
resume.json
Work
Senior engineer on the Cloud Engineering / SRE team and the team's Datadog subject-matter expert, operating cloud security, reliability, and observability for AWS and Azure workloads in a HIPAA/PCI-regulated healthcare environment.
- Datadog SME and observability lead: set monitoring, alerting, SLO, and cost-governance standards and built a developer-enablement program across five engineering domains.
- Cloud security posture program: architected CIS L1 hardening into a Golden AMI pipeline (600+ hosts), deployed AWS Inspector across the EC2/ECR fleet, and ran CSPM triage; reduced critical findings and automated PCI/HIPAA audit evidence.
- AI engineering enablement: established security and approval governance for AI-agent and Model Context Protocol (MCP) tooling and evaluated Amazon Bedrock cross-region inference.
- On-call and reliability: led the stabilization and standardization of team on-call practices and own the rotation schedules in GoAlert; improved visibility and incident management cut on-call incidents from 3–5 weekly to 1–2 monthly. Designed a DORA metrics pipeline for delivery-performance visibility.
AWSAzureDatadogTerraformKubernetesGitHub ActionsGoAlertPythonSRE
Independent practice delivering production-grade AI automation to small and medium businesses — automations, custom tools, and AI workflows on AWS Lambda, Step Functions, Bedrock, and DynamoDB, with Datadog monitoring and Terraform IaC.
- Fixed-price engagements shipped to production with monitoring, infrastructure as code, documentation, and a warranty.
- Publish open-source starter templates and practitioner content on applying AI to real operational work.
AWS LambdaStep FunctionsBedrockDynamoDBDatadogTerraformPythonCI/CD
Automated CI/CD pipelines and release management for enterprise cloud; improved uptime and monitoring; led root-cause investigations; supported the on-premise-to-Azure migration.
Automated commercial software deployments in enterprise cloud; improved performance, uptime, monitoring, and alerting; led the transition to managed Kubernetes, including container development.
Re-engineered the AWS log-aggregation platform serving 270+ business units into serverless Python, cutting on-premise load and processing cost by roughly 30%. Built front- and back-end validation tools to verify log receipt.
Set and executed technology strategy for a charter-school organization with 1,200+ devices across three campuses: network infrastructure (VLANs, DHCP, caching DNS, OpenLDAP), switch-fabric redesign, student-information-system migration (~50% cost reduction), Google Workspace for Education rollout.
Built and ran the department's instructional technology program: 100+ workstations with automated software deployment, an LDAP-based single sign-on system, and automated print-cost accounting.