A_

Anush

AI/ML engineer working across GenAI pipelines, real-time data engineering, and quantitative systems.


I build production GenAI systems — RAG pipelines, LLM-based agents, and guardrailed automation — alongside real-time data engineering and quantitative risk systems. Most recently at LTI Mindtree, with a background spanning distributed data pipelines (PySpark), LLM integration (LangChain, Claude/GPT-compatible APIs, vector search), and statistically-grounded trading systems.

SamacharX — Real-Time News Intelligence Pipeline

  • Built a low-latency Medallion Architecture (Bronze/Silver/Gold) pipeline, containerized with Docker, to ingest, validate, and summarize 500+ daily unstructured news records using PySpark for distributed real-time processing.
  • Applied LLM-based NLP (via Claude/LLM APIs) to extract and summarize actionable information from raw text, cutting manual triage while preserving fidelity for multi-channel delivery.
  • Implemented semantic deduplication using Pinecone vector search over news headline embeddings, suppressing re-publication of previously covered stories at a 0.8 cosine similarity threshold.
  • Designed a human-in-the-loop guardrail system via Telegram, ensuring 100% quality control on AI-generated output before production release.
  • PySpark
  • Docker
  • LangChain
  • Claude API
  • Pinecone
  • Telegram API

MailSaathi — LLM Document Processing Agent

  • Architected a context-aware AI agent that parses unstructured email text to extract structured action items and metadata, reducing manual sorting time by 70%.
  • Applied strict JSON schema validation and guardrails to guarantee data integrity when converting free-text LLM output into reliable structured records.
  • Established a CI/CD pipeline (GitHub Actions) for automated testing and linting, ensuring production-grade reliability and documentation.
  • Gathered requirements iteratively to tune extraction accuracy and prompt templates for real-world, high-variance document formats.
  • Python
  • GitHub Actions
  • Claude/LLM APIs
  • JSON Parsing

Quantitative Risk & Automated Trading System

  • Built prop_mm, a nine-module Python package for dynamic position sizing — combining Kelly criterion, CLT-based P&L modeling, Monte Carlo simulation, and fixed-fractional risk sizing for live capital allocation.
  • Engineered a real-time MT5 Expert Advisor (MQL5) for automated trade execution and risk management, including conditional stop-loss adjustment and an equity kill-switch.
  • Ran a rigorous post-mortem on a failed trading challenge, isolating commission drag and position-sizing errors via R-multiple analysis to derive a data-backed 0.20–0.25% fixed-fractional risk framework.
  • Backtested a mean-reversion signal across 15 liquid crypto assets (2 years of daily data), incorporating India-specific tax rules (30% VDA rate, 1% TDS) for execution-realistic strategy evaluation.
  • Python
  • MQL5
  • Monte Carlo Simulation
  • Statistical Modeling
Languages & databases
Python, SQL (MySQL)
GenAI, LLMs & agents
LangChain, RAG architecture, prompt engineering, vector databases (ChromaDB, Pinecone), LLM APIs (Claude, GPT-compatible), AI agent workflows, guardrails & system rules
Data engineering
Apache Spark, PySpark, FastAPI, Pydantic, Pandas, NumPy, Scikit-learn, XGBoost
Cloud, DevOps & MLOps
Git/GitHub, Docker, CI/CD (GitHub Actions), AWS, REST APIs / microservices (FastAPI), Linux
Quantitative systems
Monte Carlo simulation, statistical modeling (CLT), MetaTrader5/MQL5, risk sizing

Open to AI/ML engineering roles, particularly in GenAI, quant research, or fintech.
Reach me at anush.jain.official@gmail.com, or find my code on GitHub.