Hi, I'm

Nanjesh Ramesh

Senior Software Engineer | Backend, Distributed Systems & Data Engineering

I design and build distributed data platforms, backend systems, and real-time data infrastructure that power large-scale financial and operational decision making.

6+ years experience · AWS · Python · Spark · SQL · Distributed Systems

Nanjesh Ramesh

6+ yrs

Professional experience

AWS & Amex

Enterprise engineering teams

$100B+

In tracked business value supported

Batch + Real-time

Data platforms built at scale

About

I'm a Senior Software Engineer focused on distributed data platforms and backend systems, with 6+ years of experience building production systems at Amazon Web Services and American Express. My work spans distributed ETL, real-time event processing, data modeling, backend APIs, observability, and operational reliability.

At AWS, I've built systems supporting tens of thousands of business workflows and over $100B in tracked business value. I focus on turning complex, unreliable data flows into scalable, observable, and maintainable platforms that teams can trust.

I enjoy solving problems where scale, reliability, and simplicity intersect. Whether I'm designing distributed ETL pipelines, backend APIs, or data models, I aim to build systems that are easy to operate, observable in production, and maintainable as products evolve. I value strong engineering fundamentals, clear ownership, and close collaboration with product and engineering teams.

What I Build

Data Platforms

Batch and real-time pipelines, lake and warehouse architecture, curated datasets.

Distributed Processing

Spark and PySpark workloads, scalable transformations, schema handling, replay and backfill strategies.

Backend Systems

APIs, GraphQL, Lambda, Spring Boot, integration services, event-driven workflows.

Reliability

Monitoring, data quality, incident response, compliance controls, SLAs and production readiness.

Professional Experience

Software Engineer, Big Data (Data Platform & Backend Systems)

aws Amazon Web Services

May 2021 – Present

Build distributed data platforms, backend services, and real-time pipelines supporting AWS investment, finance, and partner measurement workflows.

  • Engineered data infrastructure powering dashboards that support leadership investment decisions across tens of thousands of business opportunities, representing over $100B in tracked value.
  • Architected reporting pipelines for large-scale credit and incentive programs, ensuring accurate issuance, redemption, and financial reporting at scale.
  • Built real-time and batch ingestion systems for hundreds of incentive programs, cutting reporting latency from 24 hours to realtime.
  • Designed a reusable JSON-driven ETL framework that reduced pipeline development effort by roughly 30%.
  • Developed GraphQL and Lambda-based attribution services and publishing workflows serving analytics and ML consumers.
  • Implemented CloudWatch, ELK, and compliance monitoring that cut Sev-2 escalations and audit discrepancies by about 40%.
  • Led production incident response, backfills, schema remediation, and cross-functional technical design.

Software Engineer

amex American Express

Jun 2020 – May 2021

Part of the Data Quality Management team, responsible for building and maintaining backend services that kept data accurate and reliable across high-volume financial systems.

  • Built and optimized REST APIs in Spring Boot, integrating with Couchbase and MySQL databases.
  • Migrated legacy services from Node.js to Spring Boot to improve reliability and maintainability.
  • Added automated data quality checks to Spark pipelines to catch inconsistencies before they reached downstream consumers.
  • Reviewed code, mentored new engineers, and put together onboarding documentation and training materials.

Projects

A mix of open-source projects and sanitized case studies based on professional work. Professional case studies are generalized to protect confidential information.

Quant Strategy Data Pipeline

Live demo ↗ Source ↗
PASS FAIL Fragmented Sources Standardize / Schema Contracts Partitioned Parquet Lake Quality Checks Queryable via DuckDB Incident Runbook

Problem: Each strategy team publishes analytics output in its own inconsistent format, making it hard to build one reliable, queryable dataset.

Solution: Built a standardization pipeline with enforced schema contracts (pandera) and partitioned Parquet storage, gated by SLO-style quality checks that either pass data through to querying or trigger a documented incident runbook, plus an Airflow DAG showing how it would run in production.

Impact:

  • Three inconsistent source formats unified into one schema-contracted table
  • Partition-pruned queries via DuckDB, standing in for Presto/Spark
  • Reproducible schema-drift incident with a full runbook and postmortem
  • Full pytest suite covering schema contracts and the incident path

Enterprise Data & Measurement Platform

REAL-TIME BATCH Fragmented Sources Kinesis / Lambda Ingest Scheduled Batch Ingest Spark Streaming Spark Batch ETL Redshift Curated Models Publishing APIs / Dashboards

Problem: Finance, investment, and partner teams needed fast, trustworthy visibility into large-scale operational data, but source systems were fragmented, reporting lagged up to 24 hours behind, and inconsistent validation put stakeholder trust in the numbers at risk.

Solution: Built an end-to-end AWS data platform with parallel batch and real-time processing paths (Kinesis/Lambda for streaming, scheduled ingestion for batch) converging through Spark into curated Redshift models, then served through standardized publishing APIs to dashboards and downstream consumers.

Impact:

  • Gave finance and leadership teams real-time visibility into hundreds of incentive programs, replacing a 24-hour reporting lag
  • Enabled confident, data-driven investment decisions across tens of thousands of business opportunities, representing $100B+ in tracked value
  • Delivered six mission-critical finance datasets (contracts, invoices, payments, and more) ahead of schedule with zero production issues, unblocking downstream teams early
  • Increased stakeholder trust in reported numbers by cutting production and audit issues through validation and standardized schema contracts
  • Improved the reliability of systems finance and partner teams rely on daily, reducing downstream reporting failures

Open Source Contributions

Top merged contributions first, out of 23 PRs across the data engineering ecosystem.

trinodb/trino

Merged

Fix TRIM/RTRIM on CHAR stripping real trailing content, not just padding

Traced a SQL correctness bug, where TRIM and RTRIM on CHAR values stripped real trailing content, back through eight years of git history to a leftover trimTrailingSpaces() call. A first redesign that deleted the affected overloads was caught in review as a regression under legacy coercion, so the merged fix keeps all six overloads and removes only the two stray calls, with a test covering the legacy path.

View PR #31012 →

apache/druid

Merged

Accept LONG as an alias for BIGINT in CAST expressions

Users of Druid's native engine write CAST(x AS LONG) since LONG is the native engine's name for its 64 bit integer type, but the SQL parser only recognized BIGINT and rejected the same cast with a parse error. Added LONG as a grammar level alias inside the CAST/SAFE_CAST/TRY_CAST production, producing the same type BIGINT already produces, without introducing a new column type or touching the shared DDL grammar.

View PR #20402 →

fivetran/great_expectations

Merged

Pass usedforsecurity=False on non-security md5 calls for FIPS hosts

Fixed a FIPS-compliance crash caused by unflagged MD5 usage across id generation, batch fingerprinting, and hash-based partitioning/sampling, adding usedforsecurity=False at every non-security call site without changing a single existing digest.

View PR #12099 →

fivetran/great_expectations

Merged

Pin UTF-8 encoding on filesystem-backed store reads

Fixed a locale-dependent bug where store values and the project YAML were written as UTF-8 but read back using whatever encoding the process's ambient locale resolved to, crashing on non-UTF-8 hosts (mainly Windows). Pinned encoding="utf-8" on every affected read/write path and added tests that force a non-UTF-8 locale to prove the fix.

View PR #12125 →

meltano/meltano

Merged

Support .python-version file for plugin installs

Added a fallback so Meltano picks up a project's .python-version file (the same convention used by pyenv and uv) when choosing which Python to use for plugin virtual environments, without overriding any explicit plugin or project-level setting.

View PR #10283 →

MrPowers/chispa

Merged

Add full_log option to suppress the diff table in equality errors

Added an opt-in full_log parameter to suppress the full row-by-row diff table in DataFrame equality failures, useful when comparing large DataFrames whose diffs would otherwise flood test output.

View PR #195 →
View all 23 contributions →

Skills

Languages

  • Python
  • Java
  • SQL
  • JavaScript / TypeScript

Big Data & Streaming

  • Apache Spark / PySpark
  • Hadoop
  • Hive
  • Kafka
  • Airflow

Cloud Platforms

  • AWS (S3, Redshift, Lambda, EMR)
  • Google Cloud (BigQuery)
  • Azure

Databases & Storage

  • MySQL
  • DynamoDB
  • Redshift
  • Couchbase

Backend & APIs

  • REST APIs
  • GraphQL
  • Spring Boot
  • Microservices

Data Modeling & Warehousing

  • Dimensional Modeling (Kimball / Star)
  • Data Lakes & Warehouses

BI & Visualization

  • Tableau
  • Power BI
  • ELK / Kibana

DevOps & Infrastructure

  • Docker
  • Kubernetes
  • CI/CD

Observability & Reliability

  • CloudWatch / ELK Monitoring
  • Incident Response

AI / LLM Tooling

  • LLMs (OpenAI, Anthropic)
  • RAG & Vector Databases

Education & Recognition

M.S. in Information Systems

May 2020

California State University, Los Angeles

B.E. in Computer Science

Aug 2018

Visvesvaraya Technological University

Contact

I'm open to Senior Software Engineer, Data Platform, and backend engineering opportunities where I can build scalable, high-impact systems.