$ WHOAMI

Aditya Kumar Singh

Applied AI Engineer

I build AI systems that understand how people actually speak.

I work across multilingual ASR/TTS, model fine-tuning, speech evaluation, LLM-powered pipelines and backend infrastructure — taking ideas from experimentation to production.

Multilingual Speech · Voice AI · Production ML

OPEN TO

Applied AI · Speech ML · Voice AI · ML Engineering

LOCATION

India · Remote

currently_building → multilingual speech systems
Deep LearningSpeech RecognitionSpeech Generation Model FinetuningAI ResearchStatistical Inference FastAPIPyTorchLLMs PostgreSQLDockerAWS

WHAT_I_BUILD/

I sit at the intersection of machine learning research and production engineering.

Speech AI

ASR · TTS · Language Identification · Speaker Recognition · Speech Quality Evaluation

Applied ML

Fine-tuning · Experimentation · Benchmarking · Human-in-the-loop Validation · Statistical Evaluation

LLM Systems

LLM APIs · Context-aware prompting · Linguistic edge cases · Embeddings · AI pipelines

Production Engineering

FastAPI · REST APIs · Docker · PostgreSQL · AWS · GCP · Linux

SELECTED_WORK/

A few systems I've built, evaluated and shipped.

01VOICE ARENA

VOICE AI EVALUATION & BENCHMARKING

Designed experiments and statistical evaluation methodologies for benchmarking TTS and ASR systems, formulating hypotheses, conducting significance tests, and analyzing experimental results to identify meaningful differences in model performance.

Speech EvaluationASR/TTSPythonStatistical Validation
Voice Arena ↗
02INDIC SPEECH

Language - Gender Identification & Speaker Recognition

Fine-tuned a wav2vec XLS-R model for language and gender identification across 13 Indian regional languages and built speaker/demographic classification using ECAPA voice embeddings and SVM.

wav2vec2SpeechBrainLIDSpeaker RecognitionSVM

Confidential — built @ Josh Talks

03ASR LAB

Indic ASR Fine-tuning & Evaluation

Fine-tuned Whisper, Wav2Vec2 and NVIDIA NeMo ASR models for Indic datasets, using multi-GPU PyTorch DDP for large-scale experimentation, benchmarking and hyperparameter optimization.

WhisperWav2Vec2NeMoPyTorch DDPASR
04COUNTER ASSISTANT

Voice → AI → Form Filling

An AI-powered voice interface for filling service-counter forms, combining speech recognition, backend APIs and an interactive frontend with multilingual support as the product direction.

SpeechFastAPIReactLLM APIs

Private repository

EXPERIENCE/

Mar 2026 — PresentGurugram, India

Applied AI Engineer — Josh Talks

  • Building production-grade multilingual speech AI systems for Indic ASR and TTS applications, with emphasis on inference optimization, evaluation infrastructure and scalable audio workflows.
  • Core contributor to Voice Arena, an automated speech-quality evaluation system spanning MOS, CMOS, MUSHRA, WER/CER, pronunciation verification and statistical validation.
  • Engineered context-aware LLM prompts for linguistic edge cases including lexical vs. numeric ordinals and Hindi transliteration while preserving transcript alignment.
Oct 2025 — Mar 2026Gurugram, India

AI Research Intern — Josh Talks

  • Fine-tuned Whisper, Wav2Vec2 and NVIDIA NeMo ASR models for Indic datasets using PyTorch DDP across multi-GPU environments.
  • Built automated experimentation, benchmarking, human-in-the-loop validation and research-analysis workflows.
  • Worked with torchaudio and Librosa for audio preprocessing, waveform normalization, spectrogram analysis and multilingual speech validation.
May 2025 — Sep 2025Vijayawada, India

Machine Learning Intern — TechtoGreen Drone & Robotics

  • Engineered a Pearson-correlation-driven LSTM-Attention forecasting architecture and data pipeline, reducing RMSE by 12% over baseline models.
  • Led the finalist team of the Bhashini Domain Innovation Challenge among 200+ competing teams.

TECHNICAL_STACK/

$ cat stack.txt

LANGUAGES

Python · Java

ML

PyTorch · TensorFlow · Transformers · Scikit-learn

SPEECH

Whisper · Wav2Vec2 · NeMo · torchaudio · Librosa

BACKEND

FastAPI · Django · REST APIs

DATA

PostgreSQL · MySQL · S3

INFRA

Docker · AWS · GCP · Linux · Git

GENAI

LLM APIs · ChromaDB

OPEN_TO/

Primary

Applied AI Engineer

Also

Speech ML · Voice AI · ML Engineer · AI Infrastructure

Work mode

India · Remote

Building something with speech, AI or intelligent automation?

I'd love to hear what you're working on.

© 2026 Aditya Kumar Singh · Built with too much coffee and Python.