Hi, I'm
Nanjesh Ramesh
Senior Software Engineer | Backend, Distributed Systems & Data Engineering
I design and build distributed data platforms, backend systems, and real-time data infrastructure that power large-scale financial and operational decision making.
6+ yrs
Professional experience
AWS & Amex
Enterprise engineering teams
$100B+
In tracked business value supported
Batch + Real-time
Data platforms built at scale
About
I'm a Senior Software Engineer focused on distributed data platforms and backend systems, with 6+ years of experience building production systems at Amazon Web Services and American Express. My work spans distributed ETL, real-time event processing, data modeling, backend APIs, observability, and operational reliability.
At AWS, I've built systems supporting tens of thousands of business workflows and over $100B in tracked business value. I focus on turning complex, unreliable data flows into scalable, observable, and maintainable platforms that teams can trust.
I enjoy solving problems where scale, reliability, and simplicity intersect. Whether I'm designing distributed ETL pipelines, backend APIs, or data models, I aim to build systems that are easy to operate, observable in production, and maintainable as products evolve. I value strong engineering fundamentals, clear ownership, and close collaboration with product and engineering teams.
What I Build
Data Platforms
Batch and real-time pipelines, lake and warehouse architecture, curated datasets.
Distributed Processing
Spark and PySpark workloads, scalable transformations, schema handling, replay and backfill strategies.
Backend Systems
APIs, GraphQL, Lambda, Spring Boot, integration services, event-driven workflows.
Reliability
Monitoring, data quality, incident response, compliance controls, SLAs and production readiness.
Professional Experience
Software Engineer, Big Data (Data Platform & Backend Systems)
aws Amazon Web Services
May 2021 – Present
Build distributed data platforms, backend services, and real-time pipelines supporting AWS investment, finance, and partner measurement workflows.
- Engineered data infrastructure powering dashboards that support leadership investment decisions across tens of thousands of business opportunities, representing over $100B in tracked value.
- Architected reporting pipelines for large-scale credit and incentive programs, ensuring accurate issuance, redemption, and financial reporting at scale.
- Built real-time and batch ingestion systems for hundreds of incentive programs, cutting reporting latency from 24 hours to realtime.
- Designed a reusable JSON-driven ETL framework that reduced pipeline development effort by roughly 30%.
- Developed GraphQL and Lambda-based attribution services and publishing workflows serving analytics and ML consumers.
- Implemented CloudWatch, ELK, and compliance monitoring that cut Sev-2 escalations and audit discrepancies by about 40%.
- Led production incident response, backfills, schema remediation, and cross-functional technical design.
Software Engineer
amex American Express
Jun 2020 – May 2021
Part of the Data Quality Management team, responsible for building and maintaining backend services that kept data accurate and reliable across high-volume financial systems.
- Built and optimized REST APIs in Spring Boot, integrating with Couchbase and MySQL databases.
- Migrated legacy services from Node.js to Spring Boot to improve reliability and maintainability.
- Added automated data quality checks to Spark pipelines to catch inconsistencies before they reached downstream consumers.
- Reviewed code, mentored new engineers, and put together onboarding documentation and training materials.
Projects
A mix of open-source projects and sanitized case studies based on professional work. Professional case studies are generalized to protect confidential information.
Problem: Each strategy team publishes analytics output in its own inconsistent format, making it hard to build one reliable, queryable dataset.
Solution: Built a standardization pipeline with enforced schema contracts (pandera) and partitioned Parquet storage, gated by SLO-style quality checks that either pass data through to querying or trigger a documented incident runbook, plus an Airflow DAG showing how it would run in production.
Impact:
- Three inconsistent source formats unified into one schema-contracted table
- Partition-pruned queries via DuckDB, standing in for Presto/Spark
- Reproducible schema-drift incident with a full runbook and postmortem
- Full pytest suite covering schema contracts and the incident path
Enterprise Data & Measurement Platform
Problem: Finance, investment, and partner teams needed fast, trustworthy visibility into large-scale operational data, but source systems were fragmented, reporting lagged up to 24 hours behind, and inconsistent validation put stakeholder trust in the numbers at risk.
Solution: Built an end-to-end AWS data platform with parallel batch and real-time processing paths (Kinesis/Lambda for streaming, scheduled ingestion for batch) converging through Spark into curated Redshift models, then served through standardized publishing APIs to dashboards and downstream consumers.
Impact:
- Gave finance and leadership teams real-time visibility into hundreds of incentive programs, replacing a 24-hour reporting lag
- Enabled confident, data-driven investment decisions across tens of thousands of business opportunities, representing $100B+ in tracked value
- Delivered six mission-critical finance datasets (contracts, invoices, payments, and more) ahead of schedule with zero production issues, unblocking downstream teams early
- Increased stakeholder trust in reported numbers by cutting production and audit issues through validation and standardized schema contracts
- Improved the reliability of systems finance and partner teams rely on daily, reducing downstream reporting failures
Open Source Contributions
Top merged contributions first, out of 23 PRs across the data engineering ecosystem.
- Trino
- Great Expectations
- dlt
- Meltano
- Delta Lake
- Apache Hudi
- Apache NiFi
- Apache Druid
- LangChain
- chispa
trinodb/trino
MergedFix TRIM/RTRIM on CHAR stripping real trailing content, not just padding
Traced a SQL correctness bug, where TRIM and RTRIM on
CHAR values stripped real trailing content, back through eight years of
git history to a leftover trimTrailingSpaces() call. A first redesign that
deleted the affected overloads was caught in review as a regression under legacy
coercion, so the merged fix keeps all six overloads and removes only the two stray
calls, with a test covering the legacy path.
apache/druid
MergedAccept LONG as an alias for BIGINT in CAST expressions
Users of Druid's native engine write CAST(x AS LONG) since
LONG is the native engine's name for its 64 bit integer type, but the SQL
parser only recognized BIGINT and rejected the same cast with a parse
error. Added LONG as a grammar level alias inside the
CAST/SAFE_CAST/TRY_CAST production, producing
the same type BIGINT already produces, without introducing a new column
type or touching the shared DDL grammar.
fivetran/great_expectations
MergedPass usedforsecurity=False on non-security md5 calls for FIPS hosts
Fixed a FIPS-compliance crash caused by unflagged MD5 usage across id generation,
batch fingerprinting, and hash-based partitioning/sampling, adding
usedforsecurity=False at every non-security call site without changing a
single existing digest.
fivetran/great_expectations
MergedPin UTF-8 encoding on filesystem-backed store reads
Fixed a locale-dependent bug where store values and the project YAML were written as
UTF-8 but read back using whatever encoding the process's ambient locale resolved to,
crashing on non-UTF-8 hosts (mainly Windows). Pinned encoding="utf-8" on
every affected read/write path and added tests that force a non-UTF-8 locale to prove
the fix.
meltano/meltano
MergedSupport .python-version file for plugin installs
Added a fallback so Meltano picks up a project's .python-version file
(the same convention used by pyenv and uv) when choosing which Python to use for
plugin virtual environments, without overriding any explicit plugin or project-level
setting.
MrPowers/chispa
MergedAdd full_log option to suppress the diff table in equality errors
Added an opt-in full_log parameter to suppress the full row-by-row diff
table in DataFrame equality failures, useful when comparing large DataFrames whose diffs
would otherwise flood test output.
Skills
Languages
- Python
- Java
- SQL
- JavaScript / TypeScript
Big Data & Streaming
- Apache Spark / PySpark
- Hadoop
- Hive
- Kafka
- Airflow
Cloud Platforms
- AWS (S3, Redshift, Lambda, EMR)
- Google Cloud (BigQuery)
- Azure
Databases & Storage
- MySQL
- DynamoDB
- Redshift
- Couchbase
Backend & APIs
- REST APIs
- GraphQL
- Spring Boot
- Microservices
Data Modeling & Warehousing
- Dimensional Modeling (Kimball / Star)
- Data Lakes & Warehouses
BI & Visualization
- Tableau
- Power BI
- ELK / Kibana
DevOps & Infrastructure
- Docker
- Kubernetes
- CI/CD
Observability & Reliability
- CloudWatch / ELK Monitoring
- Incident Response
AI / LLM Tooling
- LLMs (OpenAI, Anthropic)
- RAG & Vector Databases
Education & Recognition
M.S. in Information Systems
May 2020
California State University, Los Angeles
B.E. in Computer Science
Aug 2018
Visvesvaraya Technological University
- Academic Honors, Cal State LA (2018–2020): Non-Resident Tuition Fee Waiver scholarship and Special Recognition for graduate research.
- Department Topper, Computer Science (2016–2018).
Contact
I'm open to Senior Software Engineer, Data Platform, and backend engineering opportunities where I can build scalable, high-impact systems.