Site Reliability Engineer

Vivekananthan M

Build Reliable Systems.
Create Real Impact.

I design, build, and operate scalable, resilient infrastructure across AWS and GCP. From production migrations to incident response, I turn complex systems into reliable, observable, and cost-efficient platforms.

Automate
Everything
Improve
Reliability
Ensure
Security
Optimize
Costs
GCPGCPScalable. Global. Open.
AWSAWSSecure. Resilient. Innovate.
Different Clouds.
A More Reliable Tomorrow.
Systems are operational
AWSAWS
Google CloudGoogle Cloud
KubernetesKubernetes
TerraformTerraform
DockerDocker
LinuxLinux
PrometheusPrometheus
GrafanaGrafana
KafkaKafka
PostgreSQLPostgreSQL
MySQLMySQL
RedisRedis
ClickHouseClickHouse
Other toolsOthers
About Me

Engineer by trade,
reliability by obsession.

I like systems that fail quietly and recover loudly — clear alerts, fast rollbacks, and dashboards that tell the truth. My work sits at the intersection of infrastructure, automation, and incident response: turning fragile, manual processes into resilient, observable platforms that let teams ship with confidence. The best infrastructure is invisible — it just works, and when it doesn't, it tells you exactly why.

◉
Reliability over noveltyBoring, predictable systems beat clever, fragile ones.
⚙
Automate the boring stuffIf it's manual and repeated twice, it becomes a script.
◈
Blameless postmortemsEvery incident is a system lesson, not a person's fault.
▤
Docs as a first-class citizenUndocumented infrastructure is infrastructure that doesn't scale.
sre@prod ~
$ kubectl get pods -n production
NAMESTATUSRESTARTS
api-gateway-7d9f8bRunning0
worker-queue-5c6d9aRunning0
cache-proxy-3f2e1cRunning0
 
$ ./deploy.sh --env=production
✓ Build passed
✓ Tests passed (128/128)
✓ Deployed to production
✓ Health checks green
 
$ _
Experience

Where I've been
building.

3+ years operating and scaling infrastructure across fintech and enterprise environments.

Site Reliability Engineer - 1
CurrentJuspay · Bengaluru, India
February 2024 — Present · 2 yrs 8 mo
Site Reliability Engineer - TPA
Juspay · Bengaluru, India
December 2022 — February 2024 · 1 yr 3 mo
Turbo Engineer Intern — Full Stack Developer
Wipro
February 2022 — December 2022 · 11 mo
3.8+Years in SRE & DevOps
2Cloud platforms owned
99.99%Uptime maintained
10xPersonal work ethic

"Work 10x and Grow 10x."
The principle that's guided every role — take ownership beyond the job description, and the growth follows.

Top Skills
GKEGCPAWS RDSKubernetesIncident Response
Currently exploring
Service MeshCost OptimizationChaos Engineering

Education

B.E. Computer Science & Engineering
P.A. College of Engineering and Technology · 2018 – 2022

Core Competencies

What I bring to
the table.

The disciplines I focus on when designing and operating production infrastructure.

Cloud Platforms

Provisioning, IAM, networking, and cost optimization across multi-cloud environments.

AWSGCPTerraform
Container Orchestration

Designing and operating production-grade Kubernetes clusters at scale.

KubernetesEKSGKE
Observability

Building metrics, logging, and tracing pipelines that surface signal, not noise.

PrometheusGrafanaVictoriaMetrics
Data Layer

Running and tuning databases and caches under real production load.

AuroraRedisClickHouse
CI/CD & Delivery

Automating build, test, and deploy pipelines for fast, safe releases.

JenkinsGitOpsPipelines
Reliability Practices

Incident response, on-call rotations, postmortems, and SLO-driven engineering.

SLOsOn-callRunbooks
Featured Projects

Infrastructure that
makes a difference.

A few highlights from my SRE journey — solving real problems at scale.

AWSKubernetes

Production Cluster Migration

Migrated production workloads with zero downtime, including sandbox (sbx) migration and infra-switch AP usage with controlled traffic movement.

EKSNetworkingTerraformObservability
GCPKubernetes

Scalable Observability Stack

Designed and operated VictoriaMetrics, vmagent and Grafana for high-cardinality metrics across multiple clusters, enabling better visibility and faster incident resolution.

VictoriaMetricsPrometheusGrafana
AWSGCP

Database & Cache Migrations

Handled migrations from RDS to Aurora, created read replicas, optimized Aurora I/O, and managed Redis, MySQL, Scylla and ClickHouse at scale.

AuroraRDSRedisClickHouse
Sample Design

A pattern I'd
reach for.

A generic, illustrative reference architecture — not a real production system — showing how I'd typically structure a resilient, observable service.

CI/CD(build & deploy)
Image Registry(container storage)
Client
DNS / CDN(edge routing)
Load Balancer(traffic split)
Service Cluster
Region A
SvcSvcSvc
Region B
SvcSvcSvc
Primary DB(with replicas)
Cache Layer(in-memory)
Message Queue(async events)
Monitoring(metrics & alerts)
Engineering Notes

Lessons from
the trenches.

Real experiences, deep dives, and practical takeaways from operating systems at scale.

Incident Review

When STS Couldn't Reach AWS

How we identified and resolved a sudden latency increase in our EKS cluster.

Cost Optimization

Reducing AWS Costs by 40%

Right-sizing, automated scaling, and intelligent shutdowns.

Best Practices

Observability that Actually Helps

Building actionable dashboards and meaningful alerts.

Bored of reading my profile? Play a game.Play →