01About 02Expertise 03Projects 04Now 05Contact
Infrastructure Engineer · IoT platforms

I build and run cloud infrastructure for IoT platforms.

sushant@infra ~ portfolio
$cat focus.txt → kubernetes | kafka | mqtt | aws | oci
Scroll

I'm an infrastructure engineer working on smart-metering and IoT platforms for utilities. I run Kubernetes across several AWS accounts, look after the Kafka and MQTT brokers that carry device traffic, and plan migrations between AWS and OCI. I like systems that are easy to operate: alerts that mean something, deploys that only touch what changed, and upgrades with a written plan.

smart meters devices EMQX mqtt Kafka strimzi AWS OCI services eks Grafana alerts

how device data moves through the platform

~90%
less CI build time, 30 min down to 3
5
AWS accounts moved to EKS 1.35
2
clouds in production, AWS and OCI
$4.7K
yearly monitoring cost traced to one alert rule
1.5M
smart meters whose data runs through the platform
24
AWS accounts in the estate I help run
$45K
recurring yearly savings from cost work
5
engineers on the infrastructure team I lead

Technical expertise

Twelve areas I work in every week, across AWS and OCI.

01

Kubernetes

Run EKS across five AWS accounts: version upgrades, node groups, custom AMIs and launch templates.

EKSnode groupslaunch templatesAMIsIAM
02

Messaging

Operate the MQTT and Kafka brokers that carry device traffic, and move managed Kafka onto Strimzi.

KafkaStrimziMSKEMQXMQTT
03

Cloud networking

Design AWS VPCs and OCI VCNs, load balancers, route tables and hub-and-spoke ingress.

VPCOCI VCNNLBroute tablesDRG
04

Observability

Tune Grafana, CloudWatch and Alertmanager so the right alerts reach people and monitoring cost stays in check.

GrafanaCloudWatchAlertmanagerVPC Flow Logs
05

CI/CD

Build pipelines that only rebuild what changed, and deploy across accounts with shared IAM roles.

change detectioncross-accountIAM roles
06

Databases

Move large MySQL datasets between RDS and EC2 with pipelines that resume after a failure.

MySQLRDSEC2mydumperSSM
07

Autoscaling

Layer Karpenter, KEDA, HPA and VPA so bursty Kafka consumers scale on lag, request services scale on load, and nodes follow the pods.

KarpenterKEDAHPAVPAGoldilocks
08

Infrastructure as code

Terraform modules with remote state, CloudFormation for the MQTT clusters, and Helm with Argo CD for application rollout.

TerraformCloudFormationHelmArgo CDGitOps
09

FinOps

Own cost across a 24-account AWS estate: right-sizing, reserved and Savings Plan coverage, cleanups, and a monthly cost-per-meter report.

Savings PlansRIsCost ExplorerAthenatagging
10

Security & governance

Account guardrails with SCPs and tag policies, SSO permission sets in place of IAM users, and WAF rate limits on auth endpoints.

OrganizationsSCPsTag PoliciesAWS SSOWAF
11

Incident response

Carry on-call for the head-end and meter-data systems, lead the debugging, and write the postmortem so the fix outlives the incident.

on-callpostmortemsrunbooksPerformance Insights
12

Team leadership

Lead a five-person team: three DevOps engineers, a support engineer and a security engineer. Sprint planning, reviews, on-call rota and sign-off on shared production changes.

sprint planningreviewson-call rotamentoring

Projects

Recent work on the platform, with the numbers that came out of it.

02
Kubernetes

EKS fleet upgrade

Moved EKS from 1.33 to 1.35 across five AWS accounts before extended-support pricing applied, working through launch template, custom AMI and IMDS issues.

EKSEC2Auto Scaling
5AWS accounts
1.33 → 1.35EKS version
03
Messaging

MQTT broker move to OCI

Designed the move of an EMQX cluster from AWS to OCI, which meant reworking TLS termination and spreading nodes across fault domains.

EMQXMQTTOCI
AWS → OCIbroker migration
04
Streaming

Kafka on Kubernetes

Moving managed MSK clusters to Strimzi. Root-caused an outage where the operator's 30-day reconcile deleted SCRAM users created outside it.

KafkaStrimziSCRAM
MSK → Strimziin progress
05
Observability

Alerting and cost audit

Traced a $13/day CloudWatch charge to one wildcard alert rule, and found routing that sent almost every alert to a null receiver.

GrafanaCloudWatchAlertmanager
$4.7Kper year found
06
Data

RDS to EC2 migration

Moved three months of data off RDS with a checksum-verified pipeline that resumes cleanly when the 60-minute session limit kills a run.

MySQLRDSSSMmydumper
resume-safeno re-copying on restart
07
Capacity

Capacity model for 10 million meters

Built a capacity model from 14 days of production data (Kafka offsets, load-balancer logs, database load) to size a 10-million-meter rollout instead of multiplying today's numbers. Pods and Kafka partitions scale out; the database writer doesn't. Its load is 85–100% writes, so read replicas can't help. Checked against a second week of data, I/O-bound consumers drained 7–35× slower than a CPU-only model predicted.

KafkaAuroraPerformance InsightsLocust
10Mmeters modelled
~5Mserverless writer ceiling
08
Kubernetes

Layered autoscaling on EKS

Moved the head-end services from static node groups to Karpenter, with KEDA scaling Kafka consumers on lag, HPA on request services and Goldilocks-guided VPA for sizing. Some services stay pinned on purpose: a decode fleet sized one-to-one with its partitions beats one that rebalances all day.

KarpenterKEDAHPAVPA
4production clusters
1.5Mmeters served
09
Streaming

MSK and ElastiCache off managed services

Cut one production cluster over from MSK to Strimzi on KRaft and from ElastiCache to Valkey 9. I measured MirrorMaker 2 first: it was I/O-bound at about 2k messages a second, far too slow to copy history, so we switched at the live tail with a small reprocess and kept the managed services warm as the rollback.

StrimziKRaftMirrorMaker 2Valkey
MSK → StrimziKafka, on KRaft
ElastiCache → Valkeycache, on Valkey 9
10
Reliability

Zone-failover drill

Ran a full failover round trip for the head-end and meter-data workloads on two production clusters: out to a second availability zone and back. The drill found two wrong assumptions in the plan: a cluster autoscaler kept reviving the source node group, and meter-data pods were pinned to a node-group name.

EKSKarpenterCluster Autoscalerrunbooks
1a → 1c → 1around trip
2clusters drilled
11
Governance

Tagging and guardrails for 24 accounts

Designed the tag taxonomy, tag policy and SCP model for a flat 24-account organisation. Backfilled canonical tags on a production account, and kept cost tags report-only on purpose: blocking launches on a missing tag would have stopped about 44% of node launches mid-scale-up.

OrganizationsSCPsTag Policiescost allocation
56% → 100%instances fully tagged
12
FinOps

Pricing flow logs before turning them on

Measured VPC Flow Logs on two production VPCs before a rollout. Logging all traffic would grow to about $15.4K a month at 1.2 million meters. Reject-only logging costs $0.93 a month today and stays at a few dollars at that scale, because rejects are almost all internet port scans and don't grow with meter count.

VPC Flow LogsCloudWatchAthena
~1,000×cost gap, all vs reject-only
$0.93per month, reject-only

Incidents I've debugged

Production problems I traced to a root cause, and what changed afterwards.

01
PostgresOCI

Fork failures that looked like network errors

Every service on a head-end cluster kept losing its database connection with "unexpected EOF". The database host ran strict memory overcommit with no swap, so Postgres had about 140 MB left to fork new backends. The fix was one kernel setting, with no restart.

PostgreSQLLinuxOKE
140 MB → 10 GBfork headroom
02
Postgres

A restart loop caused by stuck queries

About sixty pods kept failing health checks. Day-old orphaned queries had pinned the vacuum horizon, tables bloated, and the database sat at its disk IOPS cap. I cancelled them, set statement timeouts, vacuumed, and added two missing indexes.

PostgreSQLautovacuumpg_cron
~10k → 1.5kread IOPS
194 → 0.9 mshottest query
03
KubernetesKafka

An autoscaler that starved its consumers

A Kafka consumer group was stuck in an endless rebalance. VPA had shrunk decode pods from 500m to about 53m of CPU, so each batch outran the poll timeout and offsets never committed. I pinned the resources and kept VPA as recommend-only.

VPAGoldilocksKEDAKafka
20.3Mmessages of lag
53m → 200mCPU per pod
04
MQTT

An MQTT cluster that couldn't start from zero

A boot script's cluster-join check found peers by EC2 tag and, when none of them answered, assumed the join had failed and stopped the broker. With every node down at once, no node could ever go first. I seeded one node by hand, joined the rest and restored the config from backup. The boot log, not the broker log, showed it, and it also showed the same check had quietly stopped the backups.

EMQXAuto Scalinguser dataSSM
187 bytessize of the empty backups that gave it away
05
IaC

A load balancer with zero targets

Clients couldn't reach a healthy MQTT cluster. A CloudFormation update had rewritten the Auto Scaling group's target groups and dropped one I had attached by hand weeks earlier. I put that target group in the template for both production stacks, including the one that hadn't broken yet.

CloudFormationNLBAuto Scaling
0 → 3healthy targets in 30 s
2stacks fixed
06
Aurora

An Aurora writer running out of memory

A meter-data Aurora writer was being restarted every few minutes. A table had just moved to 1,219 daily partitions, and a hot UPDATE by id didn't filter on the partition key, so each call opened every partition across about 200 connections. I diagnosed it from RDS events, Performance Insights and logs, and the change was rolled back the same day.

Aurora PostgreSQLpartitioningPerformance Insights
1,219partitions
~7,000relations opened per UPDATE
07
Data

A time-series database at 97% disk

I dropped old TimescaleDB chunks table by table, because one big drop exhausts the lock table. Then I shrank the root volume by cloning onto a smaller disk with the same filesystem UUID. I boot-tested the clone before the swap; the IPs stayed the same and downtime was about ten minutes.

TimescaleDBEBSrsyncGRUB
811 GBfreed
950 → 120 GiBroot volume

Experience

Three years at Polaris Smart Metering, from intern to leading the infrastructure team.

  • DevOps Engineer · Infrastructure Team Lead
    Polaris Smart Metering · team of 5 · 24 AWS accounts · 4 production EKS clusters
    May 2025 – now
  • AWS Cloud Engineer
    Polaris Smart Metering · IAM & SSO model · Redis to ElastiCache · Kafka to MSK · EMQX cluster for 100K+ connections
    Nov 2023 – Apr 2025
  • Cloud Engineering Intern
    Polaris Smart Metering · mapped the AWS footprint · first alarms and dashboards · fleet tagging
    Aug – Oct 2023
  • B.Tech, Computer Science
    B.K. Birla Institute of Engineering & Technology, Pilani
    2020 – 2024
01
TalkKCD Gujarat 2026

Metering a Nation on Open Source: What Happens When 250 Million Smart Meters Talk to Kubernetes

The open-source path a meter reading takes: MQTT on EMQX, Kafka on Strimzi, microservices on Kubernetes, watched by Prometheus, Grafana and Loki. And what it costs to run one meter.

EMQXStrimziKubernetesPrometheus
19 Sep 2026Ahmedabad
02
TalkKCD Gujarat 2026

The Price of a Meter: Autoscaling Kubernetes Under a Per-Unit Cost Ceiling

Layered autoscaling with KEDA, HPA, VPA and Karpenter, with cost per meter as the number that settles each scaling decision.

KEDAHPAVPAKarpenter
19 Sep 2026Ahmedabad
03
Certified

Certifications

Four current certifications across Kubernetes, AWS and Terraform.

DEAAWS Data Engineer · May 2026
CKAKubernetes Administrator · Dec 2025
TerraformHashiCorp Associate · Sep 2025
SAAAWS Solutions Architect · Jul 2025

What I'm working on

  • Moving Kafka from MSK to Strimzi
    Kubernetes · Kafka
    in progress
  • Network observability with VPC Flow Logs
    AWS · production accounts
    in progress
  • Cutting infrastructure cost per device
    Cost program
    in progress
  • Capacity planning for a 10-million-meter rollout
    Kafka · Aurora · load modelling
    in progress
  • A new state-utility platform, built as code from day one
    Terraform · Argo CD · APISIX · dev, stage and prod repos
    in progress

Let's talk

Open to infrastructure, platform and SRE roles. Email is the fastest way to reach me.

Emailnagilsushant@gmail.com
GitHubgithub.com/your-handle
LinkedInlinkedin.com/in/sushant-nagil
Available

Looking for my next infrastructure role

  • RolesInfrastructure, Platform, SRE
  • StackKubernetes, Kafka, MQTT, AWS, OCI
  • Experience3 years, leading a team of 5
  • Based inJaipur, India