
Ifthekhar Mohammed
Senior DevOps / Cloud / SRE Engineer | AWS, GCP & Self-Hosted AI
AWS · GCP · Azure · Kubernetes · GitOps · LLMs
I'm a senior DevOps, Cloud, and SRE engineer with 7+ years operating production infrastructure on AWS and GCP, and Azure when the workload requires it. I design and run Kubernetes platforms on EKS and GKE: private clusters, Terraform landing zones, workload identity, and GitOps with Argo CD and Helm. Git is the source of truth. Promoting an environment is a reviewed change, not kubectl from a laptop.
I own the path to production—GitHub Actions with OIDC, scanned images in ECR, Artifact Registry, and ACR, and promotion across Dev, QA, and Prod without long-lived keys. Observability is Prometheus, Grafana, and logs in CloudWatch or Google Cloud Logging. Security is SSO, RBAC, threat detection, and least privilege for humans and workloads. Cost sits beside reliability: rightsizing, log retention, and idle GPU or NAT spend are part of the same job.
I also run self-hosted LLMs on Kubernetes (vLLM / Ollama) on GPU node pools, behind a private OpenAI-compatible API, so prompts and documents never leave the VPC. Model versions live in Git; rollbacks are GitOps. I am open to senior DevOps, Cloud, SRE, and AI-infrastructure roles where that operating model is the work.
Experience
DevOps Engineer
August 2023 – PresentGrey Parrot Systems
Clients: Belle Tire, GE HealthCare
- Owned production Kubernetes on AWS (EKS) across Dev, QA, and Prod: Terraform-managed VPC, private API, IRSA, node groups, Argo CD, and Helm GitOps.
- Stood up the same platform pattern on GCP: GKE, Artifact Registry, Cloud Storage, Cloud IAM / Workload Identity, Cloud DNS, and Cloud Monitoring so workloads were not locked to one cloud.
- Built GitHub Actions pipelines that build, scan, and publish images to ECR and Artifact Registry, then promote environments without long-lived cloud keys (OIDC).
- Wrote Terraform modules with remote state per environment for AWS and GCP networking, clusters, and add-ons so Dev/QA/Prod stay parallel.
- Ran kube-prometheus-stack, Grafana, and Fluent Bit into CloudWatch and Google Cloud Logging; dashboards for pod health, load-balancer traffic, and node/GPU saturation.
- Hardened the platform: CrowdStrike on the cluster, GuardDuty and Inspector on AWS, least-privilege GCP IAM, SSO, and Kubernetes RBAC.
- Cut cloud spend on both AWS and GCP: Savings Plans / committed use, log retention, rightsizing node pools, and catching idle GPU and NAT cost.
- Deployed self-hosted LLMs on Kubernetes (vLLM / Ollama) on GPU node pools so internal prompts and documents stay in the VPC instead of a public model API.
- Shipped an internal OpenAI-compatible inference API in front of those models: auth, rate limits, model version in Git, GitOps rollbacks, and Grafana on latency and errors.
- Connected internal assistants / RAG to that private endpoint, with image scanning, egress limits, and no secrets in prompts or ConfigMaps.
- Partnered with app teams so promotions are a Git change—Helm values, OIDC, and environment parity—not kubectl apply from a laptop.
System Engineer — DevOps
Sept 2019 – Aug 2022Tata Consultancy Services, India
Clients: Euroports, Forth Port, Adani
- Worked with a team of 10 to migrate a product from on-premises to Kubernetes on cloud landing zones, with workloads portable across AWS, Azure, and GCP.
- Designed cloud foundations—compute, object and block storage, load balancing, and autoscaling—first on AWS (EC2, S3, EBS, ELB) using patterns that transfer to Azure VM / Blob / Load Balancer and GCP Compute / Cloud Storage.
- Progressed CI/CD from Jenkins to GitHub Actions and Azure DevOps, cutting manual deployment work by about 70% and adding tests and image builds to the path to production.
- Deployed Kubernetes for high availability, improving application reliability by about 30% with health probes, rolling updates, and environment isolation.
- Automated patch management with Ansible across Windows and Linux, reducing critical vulnerabilities by about 80%.
- Ran ELK and Prometheus for real-time logging and metrics, catching issues earlier and reducing repeat incidents by about 25%.
- Introduced Terraform for repeatable network and compute so lower environments no longer drifted from production by hand-built consoles.
- Standardized secrets and least-privilege identities for pipelines and workloads instead of shared accounts and long-lived keys.
- Supported hybrid operations: on-prem servers, cloud landing zones, and container platforms in the same release cadence.
Education
Masters in Computer Science
Machine Learning, Cloud Computing
Received Graduate Scholarship
Certifications
AWS Solution Architect Associate
Certified Kubernetes Administrator
AWS Cloud Practitioner
Google Cloud Security Fundamentals
Python Programming
Technical Skills
Cloud Platforms
CI/CD & DevOps
Container & Orchestration
AI & MLOps
Messaging & Streaming
Monitoring & Quality
Programming Languages
Frameworks
Databases
Version Control
Leadership & Achievements
Volunteer at local tech meetups
Participant in the CNCF Hackathon
Participant in the AWS re:Invent 2024
Awarded as the top performer among a team of 8