i am a master's student in computer science at arizona state university. before that, i spent 4 years building software at pwc, infosys, mindstix and browserstack.
these days most of my time goes to research. three papers accepted so far, listed under research, and a lot more time spent on the ones that did not work.
outside of all that, i’m trying to make time for the people, places, and little things that make life interesting.
experience
- developed full-stack data analysis and visualization applications using react, python, and fastapi to explore large-scale census datasets, delivering insights that support research and policy-driven decision making
- built rest web services and automated data-processing workflows, collaborating cross-functionally with researchers to translate ambiguous requirements into working product features
- designed and built event-driven microservices using nestjs, azure service bus, and cosmos db to manage homesites, milestones, and alerts across 6+ modules, improving scalability and decoupling services
- executed data migration of 1m+ records from sql to azure cosmos db using azure data factory, achieving <5 minutes downtime via staged pipelines and data validation checks
- built distributed change propagation system with job orchestration and cosmos db stored procedures, updating 10k+ homesites and ensuring consistency across microservices
- implemented job tracking and retry mechanisms for long-running operations, preventing partial failures and ensuring reliable execution across distributed services
- led team of 5 developers and 3 qa engineers delivering features using node.js and react, improving sprint predictability and ensuring consistent on-time releases
- diagnosed and resolved multiple critical production issues impacting 40 areas, caused by inconsistent state across services, restoring system stability within sla timelines and saving hours of manual data correction
- optimized etl pipeline processing 120 csv files every 2 hours by introducing stream-based parsing and improving sql queries, reducing processing time from 2 hours to under 1 hour and improving system stability
- improved backend apis using graphql by optimizing resolvers and reducing redundant data fetching, decreasing response time and improving efficiency for client-side data consumption
- refactored 10,000+ lines of legacy code, reducing production defects by 60% based on release tracking and improving maintainability for faster feature development
- built serverless workflows using azure functions for asynchronous processing, improving scalability and reducing manual intervention in data handling tasks
- awarded excellence award (top 10 out of 250+ employees) and secured 2nd place out of 40 teams in a company-wide cloud innovation hackathon
education
arizona state university, master of science, computer science
don bosco institute of technology, bachelor of engineering, computer science
selected work
projects
This website: a plain typeset page that gets damp wherever someone moves a cursor over it.
At rest it is a document. One typeface, black on off-white, hairline rules. Everything is already in the HTML, so rows open the moment you click them and every open row is a link you can send to someone.
Under the text sits a single WebGL canvas holding a low-resolution wetness buffer. Your cursor writes into it, other visitors' cursors write into it in their own colours, and it dries back to nothing in a few seconds. The chat, the postcards and the ripple all draw into that same buffer.
what I learned
I built the plain version first and put it online before touching any effects. After that, every effect had to be worth adding to a page that already worked, and a few of them were not.
Next.js, TypeScript, WebGL, PartyKit, Motion, Canvas 2D, sharp
An agent that does a maintainer's first fifteen minutes on a new GitHub issue: classify it, find duplicates, guess which files need to change, and hand a reviewer a plan to approve or reject.
Closed issues already carry the maintainer's labels, duplicate links and the commit that fixed them, so the agent is evaluated against thousands of real decisions without hand-labeling anything.
[Results: which architecture won, at what cost per issue.]
what I learned
[What I learned, including what did not work.]
Python, FastAPI, LangGraph, MCP, PostgreSQL, pgvector, React, TypeScript, OpenTelemetry
Answers analyst questions over SEC filings and federal rules with paragraph-level citations, and says so when the documents do not support an answer.
It also diffs risk disclosures year over year and pulls obligations into typed rows a reviewer can accept or reject. The filings are long, badly structured and full of boilerplate, which makes retrieval a real problem rather than a demo.
[Results: what chunking, fusion and reranking actually changed.]
what I learned
[What I learned, including what did not work.]
Python, FastAPI, LlamaIndex, PostgreSQL, pgvector, FAISS, Neo4j, React, TypeScript
A self-hosted gateway that every LLM call routes through, with the evaluation, policy and serving stack behind it.
One OpenAI-compatible endpoint gives teams per-key budgets, failover, caching, PII redaction and tracing. Requests go to a small fine-tuned model first and escalate to a frontier model only when it is unsure.
[Results: frontier spend saved at equal quality.]
what I learned
[What I learned, including what did not work.]
Python, FastAPI, PostgreSQL, Redis, vLLM, PEFT, TRL, MLflow, Kubernetes, Terraform, Grafana
A small open model turned into a domain specialist that can run inside a network that is not allowed to send documents to an outside API.
Corpus construction, continued pretraining, SFT, DPO, quantization and serving, with a scaling study to justify every GPU-hour and a comparison against retrieval alone at matched cost.
[Results: scaling efficiency, and where retrieval matched training.]
what I learned
[What I learned, including what did not work.]
PyTorch, Hugging Face, FSDP, DeepSpeed, PySpark, Databricks, vLLM, Slurm
An order and fulfillment platform where checkout touches four services and no single database transaction can span them.
Checkout runs as a saga, built both ways (choreography and orchestration) and compared under load, with compensating steps that release stock and refund payment when anything fails partway through.
[Results: which saga style won, and why.]
what I learned
[What I learned, including what did not work.]
Java, Spring Boot, Kafka, PostgreSQL, MongoDB, Redis, Elasticsearch, Debezium, Kubernetes
A real-time collaborative whiteboard: several people edit the same board at once, see each other's cursors, and lose nothing when a connection drops or a server restarts.
Edits merge as CRDTs, so there is no central sequencer, and gateway servers scale out behind a Redis backplane. Offline edits merge cleanly on reconnect.
[Results: fan-out latency as clients and servers grow.]
what I learned
[What I learned, including what did not work.]
Go, WebSockets, CRDT, Redis, MongoDB, DynamoDB, PostgreSQL, React, TypeScript
research
SNIPER-SQL ranks #1 globally on the BIRD-CRITIC-1.0-Open leaderboard (570 multi-dialect tasks) with a database-free validate-and-repair pipeline: dialect-strict chain-of-thought generation, adaptive decoding, and a static schema-aware SQL validator.
It generalises across six large language models from five model families, spanning 17B to 235B parameters, and raises SQL repair success for both frontier and open-weight models, with paired significance testing.
Further details and the paper will be released once the paper is accepted.
Literature-grounded question answering requires a system to retrieve the right paper, locate the right evidence within it, and generate an answer faithful to that evidence. These are three separate stages, each with its own failure mode. We present a stage-wise diagnostic study of this pipeline on LitTraceQA, a shared task requiring paper retrieval, coarse evidence localization (table/figure/page), and structured answer generation. Our pipeline combines cross-model hypothesis diversity for retrieval, a filter-then-rank evidence localizer (caption-pattern match, embedding similarity, evidence-type gate), and cross-model semantic entropy for generation-time answer aggregation and confidence estimation. We use an oracle-substitution ablation ladder, successively handing the pipeline the gold paper and then the gold evidence value, and reading off the resulting change in each stage’s own metric. These are sequential oracle-intervention gains, not an additive error decomposition: the stages are measured with different quantities and are not independent. Read that way, the ladder points at evidence localization as the dominant remediable bottleneck — multiple-choice accuracy rises from 0.42 with predicted evidence to 0.76 with the gold evidence value, and end-to-end accuracy stays near 0.29 even on the questions where retrieval already returned every gold paper. We also test whether cheap uncertainty signals computed at retrieval time predict which stage will fail before it does. Along the way we surface two infrastructure-level findings with direct implications for anyone building literature-grounded QA systems. First, evidence page numbers in this and similar benchmarks are keyed to a paper’s canonical published version rather than its arXiv preprint: a real failure mode when a paper has been substantively revised, though a systematic paired ablation across 45 questions finds it is an occasional rather than pervasive hazard in this pool. Second, a substantial fraction of papers hosted on OpenReview (ICLR/ICML/NeurIPS) are unreachable by ordinary scripted HTTP requests due to a JS bot-challenge, a venue-correlated failure mode unrelated to any modeling choice. We report our diagnostic protocol, quantitative results with bootstrap confidence intervals given the benchmark’s small public dev set, and a qualitative failure taxonomy, framing this as a resource and lessons-learned contribution rather than a leaderboard claim. All multiple-choice numbers are a best-case diagnostic condition: they assume access to the gold option text, which a genuine blind submission would not have.
Dense retrieval embeds each query as a single vector, inheriting the failure modes of whichever language model generates the representation. HyDE partially addresses this by embedding a generated hypothetical answer rather than the raw question, but the single-generator design is brittle: a confident hallucination deflects the query vector away from relevant passages, and on multi-hop benchmarks four of the five individual HyDE variants we tested degrade retrieval below the direct-query baseline (the fifth is roughly at parity).
We propose Cross-Model Hypothesis Aggregation (CMHA), which queries N diverse LLMs from distinct model families, embeds each generated hypothesis, and uses the normalized centroid as the final query vector. Per-model errors are geometrically uncorrelated across families, so centroid averaging cancels idiosyncratic mistakes while reinforcing shared factual content. On BEIR benchmarks with BAAI/bge-large-en-v1.5, CMHA with N=4 achieves Recall@10 of 0.8242 on Natural Questions (+8.9pp over direct query) and 0.8166 on HotPotQA, recovering a 7.7pp single-model HyDE degradation; pairing cross-model diversity with K=3 temperature samples per model further improves peak Recall@10 to 0.8313 on NQ and 0.8328 on HotPotQA. All headline gains are statistically robust (bootstrap 95% CIs excluding zero) and hold, with a larger effect size, under a stricter joint-recall metric. We also find, to our knowledge for the first time, that embedding the raw output of chain-of-thought thinking models catastrophically breaks HyDE-style retrieval by up to 68.2pp: unparsed reasoning traces are not document-like text and embed far from the relevant passage region. A diversity-based adaptive routing policy achieves 99.9% of full CMHA recall at only 3.8 average LLM calls per query, making the approach practical for latency-sensitive vector database deployments.
Energy poverty is nearly absent from NLP-for-social-good, and the little existing work is either static retrieval/QA or relies on carbon-intensive cloud LLMs, a self-defeating "computational irony" for a humanitarian setting. We present EqGrid, a closed-loop simulation in which a low-frequency, open-weight LLM policy agent sets price and carbon bounds and targeted subsidies over a community of empirically-grounded household personas, while high-frequency multi-agent RL traders clear a continuous double auction constrained by a physical distribution grid (IEEE-33-bus with Dynamic Operating Envelopes). Our contribution is threefold and directly addresses how to measure the social impact of AI: (i) grounded personas (region-matched socio-demographics) whose load curves are checked for shape and level realism against real smart-meter data; (ii) formal energy-poverty equity metrics (Energy Burden, Gini of EB, LIHC) showing the intervention reduces burden inequality without raising net grid cost; and (iii) a compute-efficiency frontier that measures how much equity performance survives compressing the policy agent from a 235B teacher down to a sub-1B model deployable on a laptop, in estimated energy/carbon per decision. A decoupled-safety design (the LLM sets bounds; a validate-and-project grid gate executes) yields zero grid-constraint violations versus 55 under direct LLM control. On energy-poverty equity, the LLM policy lowers the Gini of energy burden to 0.305 (from 0.351) and mean burden by 28% while cutting cost (outperforming a tuned rule baseline), and a 3B-active model retains 95% of the benefit at roughly 9x lower inference energy than the teacher, with even a 0.8B on-device model retaining 92% at roughly 24x lower energy. We will release code and configs.