Skip to content
View Sparsh-Kumar's full-sized avatar

Block or report Sparsh-Kumar

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Sparsh-Kumar/README.md

Sparsh Kumar

Senior Software Engineer · Backend Systems · Distributed Platforms · Data & AI Infrastructure

Software Engineer with 7+ years of experience designing backend services, data platforms, distributed workflows, and production-ready AI applications.

Backend Architecture Data Platforms Distributed Systems Applied AI

GitHub profile LinkedIn profile Email Sparsh Kumar Call Sparsh Kumar


Engineering Profile

I design backend systems from service boundaries and data contracts through deployment and operation. My work spans asynchronous services, APIs, ingestion pipelines, messaging infrastructure, storage abstractions, and LLM-backed workflows.

I focus on the tradeoffs that keep systems useful as they grow. Reliability, scalability, maintainability, delivery speed, and operational clarity are treated as design inputs rather than cleanup work.

Focus Applied work
System architecture Independently deployable services, reusable libraries, stable APIs, and explicit ownership boundaries
Reliability Acknowledgements, recovery paths, deduplication, schema validation, migrations, and layered tests
Data platforms Real-time ingestion, relational and document storage, columnar data, and analytical query engines
Delivery Locked dependencies, container builds, infrastructure as code, CI/CD, and environment separation

Selected Engineering Work

In progress

Distributed data platform

Python · WebSockets · Docker · PostgreSQL · MongoDB · Terraform

Building a data platform that isolates ticker, trade, and order book ingestion into independently deployable Python services. Each service owns its dependencies, lint checks, tests, and container build. Shared libraries provide consistent logging, exception handling, and MongoDB or PostgreSQL access. Separate Compose stacks and Terraform environments keep local and production concerns explicit while allowing new ingestion protocols to follow the same operating model.


Backend AI system

Python · Flask · MongoDB · OpenAI · RSS · Groww API

A scheduled backend system that combines external news feeds with portfolio data and structured model output. The pipeline deduplicates inputs, chooses the correct processing mode, validates responses before persistence, and preserves an auditable decision history in MongoDB. Versioned migrations, guarded behavior when live data is unavailable, Flask APIs, a responsive dashboard, and layered tests support the complete workflow from ingestion to presentation.


Messaging infrastructure

TypeScript · Node.js · Redis Streams

A message queue built directly on Redis Streams to make asynchronous delivery behavior explicit. Producer and blocking consumer APIs use consumer groups, acknowledgements, and pending-message recovery. The implementation exposes work distribution, redelivery, persistence choices, and horizontal scaling without hiding the underlying broker semantics behind a framework.


Extensible data library

TypeScript · Node.js · Web Data Extraction

A TypeScript library that turns multiple categories of financial information into a consistent programmatic interface. Its asynchronous API covers company statements, ratios, ownership data, ratings, announcements, and reports. Source-specific extraction remains behind adapters, which isolates provider behavior and keeps the public contract stable as integrations are added.


RAG data pipeline

Python · Haystack · OpenAI · PDF Processing

A Python pipeline that converts unstructured PDFs into cleaned transcript collections and retrieves relevant passages for downstream analysis. Document preparation remains separate from Haystack and GPT-4o-mini retrieval, which allows the corpus to be rebuilt independently and keeps ingestion concerns outside the model-facing workflow.

Technical Toolkit

Area Technologies
Languages Python, TypeScript, JavaScript, SQL
Backend & APIs Node.js, Flask, NestJS, REST, WebSockets
Distributed Systems Kafka, Redis Streams, consumer groups, asynchronous processing
Data Platforms PySpark, Parquet, Apache Iceberg, Trino, real-time ingestion
Databases PostgreSQL, MongoDB, Redis
Cloud & Infrastructure AWS, GCP, Docker, Kubernetes, Terraform
Applied AI OpenAI APIs, RAG, Haystack, LangChain, LangGraph
Engineering Practice System design, pytest, Ruff, GitHub Actions, CircleCI

Current Focus

  • Designing reliable distributed services with clear ownership boundaries
  • Building data and AI platforms from ingestion through APIs and user-facing workflows
  • Evaluating tradeoffs across consistency, availability, latency, complexity, and cost
  • Improving operational readiness through tests, automation, observability, and failure recovery

Interested in senior backend, platform, distributed systems, and data infrastructure roles.

GitHub · LinkedIn · Email · Phone

Popular repositories Loading

  1. Docker-Learning-Notes Docker-Learning-Notes Public

    A repository for interview preparation on the topic of docker with explanation of all commands and easy description.

    JavaScript 3

  2. Highlevel-Hiring-Challenge-Backend Highlevel-Hiring-Challenge-Backend Public

    This is an appointment booking API for Dr. John, which is part of the Highlevel Hiring project.

    TypeScript

  3. Highlevel-Hiring-Challenge-Frontend Highlevel-Hiring-Challenge-Frontend Public

    This is an appointment booking API for Dr. John, which is part of the Highlevel Hiring project.

    TypeScript

  4. Redis-Streams-Message-Queue Redis-Streams-Message-Queue Public

    Implementation of message queue using Redis as message broker from scratch.

    TypeScript

  5. Finsight Finsight Public

    A TypeScript/Node.js library to fetch comprehensive stock-related information such as financials, reports, shareholder patterns, credit ratings, sentiments, and more from multiple sources.

    TypeScript

  6. Quarterly-Concall-Evaluation-Implementation Quarterly-Concall-Evaluation-Implementation Public

    This project automates the extraction and cleaning of stock concall transcripts from PDFs, preparing them for analysis. It leverages a RAG pipeline with Haystack and GPT models to summarize key poi…

    Jupyter Notebook