ALL SYSTEMS OPERATIONAL— AVAILABLE FOR HIRE
VIKAS PALAKURTHI · SENIOR SRE / DEVOPS / OBSERVABILITY · AUSTIN, TX

I keep production boring.

Senior SRE / DevOps / Observability engineer for the platforms everyone else depends on — large-scale Elasticsearch logging, Kafka pipelines, Kubernetes fleets, and alerting that actually means something. When systems go quiet, I ship products of my own.

Vikas Palakurthi
VIKAS PALAKURTHI · SENIOR SRE / DEVOPS
$ identity --verify
✓ vikas.palakurthi — human, confirmed
✓ region: austin-tx · open_to_work
$ ▍
ELASTICSEARCH·KAFKA·KUBERNETES·EKS·HELM·PROMETHEUS·GRAFANA·OPENTELEMETRY·TERRAFORM·ARGOCD·CLAUDE API·MCP·PAGERDUTY·PYTHON·GO·FASTAPI·REACT·NEXT.JS·
ELASTICSEARCH·KAFKA·KUBERNETES·EKS·HELM·PROMETHEUS·GRAFANA·OPENTELEMETRY·TERRAFORM·ARGOCD·CLAUDE API·MCP·PAGERDUTY·PYTHON·GO·FASTAPI·REACT·NEXT.JS·
10y
running production infrastructure — AWS, Kubernetes, and everything observability
15 TB/day
active Elasticsearch ingestion wrangled across stg/prod clusters
2,000+
dashboards & alert rules migrated Datadog → Prometheus at Apple
0
visible panic events during incidents (externally, anyway)
$ whoami

Reliability engineer by trade. Builder by habit.

I build and run the platforms other engineers depend on: large-scale Elasticsearch logging clusters, Kafka streaming pipelines, Kubernetes fleets on EKS, and the alerting that ties it all together. Most recently I did exactly that at Apple, leading a team of 10 through an enterprise telemetry migration.

On the side I design, build, and operate my own products end to end — a premarket options analytics engine and an AI-powered trading journal among them. Shipping solo teaches the things on-call can't: scope, product sense, and owning every layer of the stack.

Now looking for my next SRE / DevOps / Observability home — Austin or remote.

Off the clock: shipping my own side projects end to end — and reading other people's postmortems so I don't star in my own.
# spec sheet
role        = "Senior SRE / DevOps / Observability"
base        = "Austin, TX"
mode        = ["on-site", "hybrid", "remote"]
core_stack  = ["Prometheus", "ELK", "Kafka", "K8s"]
side_quests = "ships own products"
status      = "interviewing"
prod — tail -f /var/log/career.log
[08:12:04] INFO cluster status: green
[08:12:07] INFO kafka lag: 0ms · groups healthy
[08:12:11] WARN coffee level below threshold
[08:12:12] INFO auto-remediation: refill ✓
[08:12:19] INFO deploy → canary 100% healthy
[08:12:23] INFO pagerduty: suspiciously quiet
[08:12:41] EVENT hiring_signal: recruiter_view
[08:12:42] INFO available=true · austin|remote
[08:12:04] INFO cluster status: green
[08:12:07] INFO kafka lag: 0ms · groups healthy
[08:12:11] WARN coffee level below threshold
[08:12:12] INFO auto-remediation: refill ✓
[08:12:19] INFO deploy → canary 100% healthy
[08:12:23] INFO pagerduty: suspiciously quiet
[08:12:41] EVENT hiring_signal: recruiter_view
[08:12:42] INFO available=true · austin|remote
$ top -o expertise

Instrumented skills.

CLOUD & PLATFORM
AWSKubernetesEKSDockerHelmIstio/Envoy
state: self-healing
IAC, CI/CD & GITOPS
TerraformCloudFormation/CDKAnsibleJenkinsGitHub ActionsArgoCD
state: zero drift
OBSERVABILITY & LOGGING
Prometheus ↗Grafana (LGTM) ↗Elasticsearch / ELKOpenTelemetryDatadog ↗Elastic APM
state: battle-tested
STREAMING & EVENTS
Kafka (Confluent)RabbitMQAWS EventBridgeStreaming pipelines
state: zero lag
RELIABILITY & INCIDENT RESPONSE
SLIs/SLOsError budgetsPagerDutyOpsgenieRunbooksPostmortemsDR failovers
state: calm under fire
SECURITY & ACCESS
VaultIAMKMSSAML SSOSSL/TLSKafka RBAC
state: least privilege
AI & LLM AUTOMATION
state: context-aware
LANGUAGES & FRAMEWORKS
PythonGoBashTypeScriptReactNext.jsFastAPI
state: always compiling
$ systemctl status side-projects --all

Running services.

Products I designed, built, and operate end to end — real users, real on-call, and no one to page but myself.

oifetcher
active
oifetcher — premarket view
oifetcher screenshot

Premarket options analytics with paying subscribers. Maps open-interest walls and key levels before the bell — and I own everything behind it: infrastructure, deploys, billing, support, and incidents.

TypeScriptReactPython
tradenarrate
active
tradenarrate — journal view
tradenarrate screenshot

AI-powered trading journal with behavioral coaching, built on the Claude API — the same LLM stack I used to automate enterprise migration at Apple, pointed at a consumer product.

Next.jsTypeScriptClaude API
$ tail -f /var/log/career.log

The log so far.

2025.08 → 2026.09
INFOApple
Led a team of 10 migrating enterprise telemetry from Datadog to Apple's internal Prometheus platform (MOSAIC): 1,000+ dashboards and 1,000+ alert rules with zero disruption to production monitoring. Built an LLM-driven transformation pipeline on the Claude Code API and MCP that cut manual migration effort by 70–80%, and moved every observability asset into GitOps with CI/CD-deployed dashboards, alerts, and recording rules.
read the deep dive →
2021.06 → 2025.08
INFOFreddie Mac
Ran the enterprise logging and monitoring platform: migrated self-managed ELK on EKS to Elastic Cloud with zero data loss, built Prometheus federation and telemetry pipelines, defined SLIs/SLOs with error budgets, carried the on-call rotation with documented runbooks, and executed yearly cross-region DR failover exercises.
2019.06 → 2021.05
INFOT-Mobile
Administered Elasticsearch and Confluent Kafka clusters across all environments — RBAC and SSL/SASL hardening, partition rebalancing and broker operations, plus custom Prometheus exporters for consumer lag and replication health.
WARNolder entries truncated (Capital One, T-Mobile · 2016–2019) — request the full resume for complete history
$ ls -la ~/credentials/

Paper trail.

Degrees and certifications, dated honestly — the current credential is the production record above.

EDUCATION
M.S. Computer Science
Texas A&M University — Corpus Christi
2016 · GPA 3.8
CERTIFICATION HISTORY
verify AWS codes ↗
  • AWS Certified SysOps Administrator — Associate2018 → 2021N1W4HNT2LF111P9J
  • AWS Certified Developer — Associate2018 → 2021D4Q4XQGKLFV4QEW2
  • AWS Certified Cloud Practitioner
  • Docker Certified Associate ↗2018 → 202012258973
$ ssh vikas@your-infrastructure

Let's keep something running together.

Austin, TX · on-site, hybrid, or remote · Senior SRE / DevOps / Observability

avg response time: < 24h · enthusiasm uptime: 100%
Vikas
— written, designed & kept online by an actual human