Summary
Senior cloud and platform engineer with 14+ years professional experience in technology, including 3 years as an AWS Senior Technical Account Manager advising enterprise customers on cloud architecture, operational excellence, resiliency, and cost optimization. Since AWS, expanded hands-on engineering depth through Senior SRE and Platform Engineering roles, designing and operating AWS and Kubernetes platforms at scale.
Self-developed an AI-powered operations and FinOps automation agent at PayPay, driving over $100K in annual cloud savings and adopted by 100+ engineers. Its success led to a secondment to deploy it at PayPay Bank as well. Deep expertise in AWS, Kubernetes, observability, infrastructure as code, and AI-powered operations.
Experience
- Self-developed and deployed CloudOps Agent, an AI-powered operations and FinOps automation platform on EKS using the Strands Agents SDK with 30+ subagents connected via the A2A protocol. Drove over $100K/yr in savings across 100+ AWS accounts spanning multiple organizations. Currently used by 100+ engineers weekly and handling 1,300+ queries per week with full observability and traceability via OpenTelemetry
- Implemented a centralized AI gateway using agentgateway to provide unified observability and spend governance across Amazon Bedrock, OpenAI API, and GCP LLM usage, introducing quotas and automated limit-increase handling that closed a gap previously contributing to runaway AI spend
- Built Savings Plans and Reserved Instance dashboards with renewal alerts, CUDOS reporting, and cost anomaly detection, all automated and integrated directly into the CloudOps Agent and Slack
- Architected the AWS infrastructure for a multi-tenant observability platform serving several PayPay Group Companies, built on VictoriaMetrics, OpenTelemetry, ClickHouse, and Grafana running on dedicated multi-region EKS clusters
- Orchestrated mission-critical workload migration from ECS Fargate to EKS with zero downtime, increasing system reliability while preserving all customer functionality
- Reduced AWS infrastructure costs by over 50% within first year through strategic resource optimization, Savings Plans purchasing, and architecture refinements
- Strengthened application reliability by implementing Helm-based deployments with robust rollback capabilities, streamlining CI/CD pipelines, and enhancing observability through Prometheus, Grafana and distributed tracing with Datadog
- Implemented comprehensive cost observability including Amazon CUDOS and custom pipelines into Holistics, providing leadership with cost visibility dashboards, anomaly detection, and proactive budget alerts
- Advised enterprise customers as a senior technical partner and voice of the customer, influencing AWS service teams and product roadmaps for services including Amazon EKS and AWS Local Zones
- Applied the AWS Well-Architected Framework, with extra emphasis on the Cost Optimization pillar, to guide enterprise customers toward more efficient, well-architected AWS environments
- Guided customers through capacity planning and workload right-sizing using end-to-end user-flow profiles, balancing cost, performance, and scalability requirements
- Partnered with enterprise cloud operations teams to resolve complex AWS issues, accelerate support escalations, and identify opportunities to improve service reliability, performance, and efficiency
- Engineered an internal automation platform using PHP7 (Laravel), JavaScript, and MySQL that reduced manual processes across various departments
- Designed and implemented CI/CD pipelines with Atlassian Bamboo, increasing deployment frequency while reducing errors
- Administered self-hosted Atlassian ecosystem (Jira, Confluence, Bamboo, Bitbucket, HipChat) on VMware infrastructure, including server maintenance, application upgrades, and MySQL database optimization
- Performed international system integration for C4ISR environments in Saudi Arabia, completing deployments 15% ahead of schedule by helping develop automation scripts that reduced system provisioning time from days to hours
- Earned achievement award for completing System Level Use Case 200+ hours under budget while exceeding all customer requirements
- Implemented and managed modern CI/CD pipelines with Jenkins and Git, replacing legacy ClearCase systems and reducing build times and deployment complexity
- Provided comprehensive IT support including desktop troubleshooting, SOP creation, SharePoint optimization, OS deployments, and hardware management
- Managed end-to-end support for diverse technology stack including desktops, printers, telecom equipment, servers, networks, surveillance systems, and mobile devices across multiple facilities
- Assisted in the daily operations of the Information Systems department working both independently and in cohesive teams
- Took on several large projects including the installation of roughly 30 wireless access points, hard wiring a newly constructed building with category 6 cables, and the installation of server hardware and software
- Provided an additional communication channel between the IS manager and other department staff working closely with the CIO
Skills & Proficiencies
Proficient
Cost Engineering & FinOps Automation (AWS Cost Explorer, CUDOS, Savings Plans, Reserved Instances, AI-driven Anomaly Detection), AI/ML Agent Systems & Governance (Strands SDK, A2A Protocol, Amazon Bedrock, agentgateway), Amazon Web Services (AWS), Google Cloud Platform (GCP), Kubernetes, Python, Infrastructure as Code (Terraform, CDK, CloudFormation), Observability (VictoriaMetrics, Grafana, ClickHouse, OpenTelemetry, Prometheus), DevOps, Linux, SQL (Postgres, MySQL, Athena), CI/CD (GitHub Actions, Jenkins, ArgoCD)
Projects
This project attempts to cut through AWS announcement noise by analyzing actual AWS usage through Cost Explorer data, fetching recent AWS announcements, and using Amazon Bedrock (Default: Amazon Nova Lite) to determine which announcements are relevant to your services. Notifications can be viewed in the CLI or sent directly to a Slack channel.
The k8s-autoscaler-benchmarker can be a useful tool for administrators and developers looking to optimize the scaling capabilities of their EKS clusters. The tool offers a streamlined process for benchmarking the performance of Karpenter and Cluster Autoscaler for EKS workloads.
This resume site deployed on CloudFront using the AWS CDK. Entirely serverless, leveraging S3, CloudFront with Origin Access Control, ACM, and Route53. Replicated worldwide via CloudFront's globally distributed edge networks.