Tazapay · AI & Analytics Intern
Cross-border payments start-up, backed by Sequoia Capital & Circle Ventures · Singapore
Hi! I'm currently an MSc student at the NUS School of Computing, working at the intersection of causal inference, machine learning, and AI safety.
This year, I've received the BlueDot AI Safety Grant (Chain-of-Thought Interpretability), and won Google Cloud Rapid Agent Hackathon (Gold), ACM UbiComp/ISWC (Gold), AKBC @ EMNLP (Gold), WMT @ EMNLP, ArabicNLP @ EMNLP (Gold), the CAR-Bench Challenge @ IJCAI-ECAI 2026 (Silver), the RecSys-HR 2026 WorkRB Challenge (Gold), AlphaNova Quant Competition (9th out of 900). Previously, I was a winner of the National Mathematics Olympiad in Russia, placing in the top 0.015% nationally.
I've also finished the University of Oxford's 'ML Representation Learning & Generative AI' Bootcamp and the Beijing University of Posts & Telecommunications (BUPT) 'Agentic AI' Bootcamp, where I placed 1st in the final agentic AI competition.
Currently, I'm working as an AI Intern (Agents & Evals) at Tazapay, a stablecoin-based cross-border payments company, and as a Research Assistant on recommender systems at HONOR (China). Alongside this, I'm completing the TARA (ARENA) and BlueDot Technical AI Safety programmes.
My previous experience includes quantitative research at Orbuc, a London-based boutique crypto trading and research firm, as well as independent work in causal inference, Bayesian experimentation, and ML systems.
A critique of the public evaluation metrics for frontier AI models, using the launch week of Claude Fable 5.1 and GPT-6 Astra as the running case.
Read the essay →A Bayesian causal-mediation test of whether an LLM's stated reasoning drives its answer or is post-hoc, with calibrated uncertainty for scalable oversight.
0.7060 macro F1 in the AKBC 2026 Shared Task on predicting complete knowledge base entries from language models. Winning system paper at the AKBC workshop, EMNLP 2026, Budapest (Sep '26).
1st place ($5k prize), 15,000 teams, for AutoSRE, an autonomous on-call engineer that triages and resolves production incidents (Jul '26).
Awarded a BlueDot Impact Rapid Grant (Jun '26) to fund Bayesian Causal Faithfulness for LLM Chain-of-Thought: calibrated uncertainty for whether a model's reasoning is faithful or post-hoc.
Winner of the Russian National Mathematics Olympiad (Moscow Institute of Physics & Technology), top 0.015% nationally.
Graduated with First-Class Honours (highest distinction) from Bayes Business School; top of cohort in Quantitative Methods & Analytics, AI & Big Data, Capstone Project, and ESADE Mergers & Acquisitions.
Full Scholarships at the National University of Singapore (Spring 2025) and ESADE Business School (Autumn 2024) exchanges, plus a fully-funded scholarship for the Beijing University of Posts & Telecommunications Agentic AI Bootcamp (Jun to Jul 2026).
Cross-border payments start-up, backed by Sequoia Capital & Circle Ventures · Singapore
Statistical Modeling & Time-Series Research · London, UK
Bayes Business School Capstone · Highest grade in cohort · London, UK
Education & Admissions Consultancy · London, UK
Brazilian e-commerce panel · 97k orders · hierarchical Bayesian causal inference in PyMC 5
Experiment-safety auditing tool · SRM detection & causal inference · optional Claude tool-use agent
Energy-utility SME churn · 14,606 customers · cost-sensitive, decision-aware modelling
Sovereign-default prediction · 34-year cross-country macro panel · 5-model benchmark + PPO from scratch
Retail credit-risk modelling · probability-of-default under an asymmetric cost matrix · Q-learning
Extending the WMT26 QEbreak findings (EMNLP 2026) beyond the challenge set
The QEbreak paper measured one thing: learned quality metrics punish added text far more than missing text. This write-up asks the wider question of what else the current LLM-based evaluators get wrong about translation quality, and why the same blind spots keep showing up across systems.
OngoingA critique of the public evaluation metrics for frontier AI models
Arena rankings, benchmark scores and lab-reported numbers get read as measurements of capability. This essay asks what each of them measures and under which conditions a score says something about capability, using the launch week of Claude Fable 5.1 and GPT-6 Astra as the running case, with the same checks applied to both labs.
Read the essay →The most overhyped threshold in statistics
Like a lot of people, when I was first introduced to the 0.05 significance threshold, I just took it for granted. Later, as I learned more statistics, I realized how confusing and misleading that hard line actually is. This essay explores why: a courtroom retelling where the p-value is the evidence, 'significant' is the verdict, and 0.05 is a fixed sentence nobody ever justified, with interactive figures for the tail area, the false-discovery rate, the dance of the p-values, and the Type I/II tradeoff.
Read the essay →A geometric reading of Hidden Markov Models & the EM algorithm
When I was learning hidden Markov models, I couldn't find an explanation that really showed how they work. This piece builds a geometric, intuitive picture of hidden Markov models, along with the EM steps, forward-backward, Baum-Welch, and the Viterbi algorithm, and even how they link to PCA. Interactive diagrams and animations throughout aim to make the picture stick in your head.
Read the essay →MSc in NUS School of Computing, specialising in Statistics
HONOR x NUS Joint Research Project, Recommender SystemsAug 2026 to Present
Research Assistant · Advisor: Prof. James Pang, NUS Business Analytics Centre
Building generative retrieval (HSTU, TIGER semantic IDs) and a multi-objective MMoE ranker for an industry feed recommender.
Bootcamps:
BSc International Business (Hons), Specialised in AI and Quantitative Methods; First-Class Honours (Highest Distinction)
Top of cohort in Quantitative Methods & Analytics, AI & Big Data, Capstone Project, and ESADE Mergers & Acquisitions.
Extracurricular Quantitative Coursework: Stanford CS229 Machine Learning, Stanford CS230 Deep Learning, MIT RES.6-012 Introduction to Probability, Stanford EE178 Probabilistic Systems Analysis, Imperial College Mathematics for Machine Learning, IBM Applied Data Science Specialization (Databases & SQL, Visualization, Python), Statistical Rethinking 2026 (R. McElreath, Max Planck).
Final grade A* · Mathematics 87%, Highest Distinction
Causal Inference (DiD, RCT design, causal DAGs, wait-list controlled trials, synthetic controls, propensity-score matching, Rosenbaum sensitivity bounds, uplift modelling), A/B Testing & Experimentation (variant design, power analysis, SRM detection, CUPED variance reduction, multiple-testing correction), Bayesian Statistics (PyMC 5, NUTS, hierarchical models, posterior-predictive checks, PSIS-LOO), Survival Analysis (Cox PH, Kaplan-Meier, Random Survival Forest), Hypothesis Testing (t-test, Wilcoxon signed-rank, Mann-Whitney U, chi-square, Fisher z), Stochastic Processes & Sequential Modeling (Hidden Markov Models, time-series, state-space methods), SHAP, permutation importance, Brier, isotonic and LOOCV calibration, bootstrap & Hodges-Lehmann confidence intervals, Effect-size estimation (Cohen's d, rank-biserial).
Supervised/Unsupervised Learning, Gradient Boosting (XGBoost, LightGBM), Anomaly & Rare-Event Detection, Neural Networks (CNNs, RNNs, LSTMs, Transformers, Two-Tower), NLP, Computer Vision, Recommendation Systems (implicit feedback, negative sampling, embeddings & vector search), Reinforcement Learning (Q-learning, DQN, PPO + GAE, Actor-Critic), Probabilistic Graphical Models, LLMs, Prompt Engineering, RAG, OpenAI API, Anthropic Claude tool-use, Zod structured outputs, LLM-as-judge evaluation, LLM Fine-Tuning, RLHF, tool-using LLM agents (LangChain, LlamaIndex), Generative AI Applications.
Python (NumPy, pandas, scikit-learn, TensorFlow, PyTorch, statsmodels, SciPy, XGBoost, LightGBM, PyMC), TypeScript, C++, R, SQL, BigQuery, Spark/PySpark, DuckDB, React 19, Vite, Vercel Serverless, Supabase (Postgres).
Google Cloud Platform (BigQuery, Vertex AI), Docker, Kubernetes, MLOps (CI/CD, Model Deployment, Feature Pipelines, ETL/Airflow), GitHub Actions, pytest, Vitest, ruff, Git/GitHub, Jupyter, LaTeX, Tableau, Power BI, matplotlib, seaborn, Plotly.
English (Fluent), Russian (Native), Ukrainian (Native), Belarusian (Fluent), Spanish (Professional Working; advanced certification, 2026).
Stanford CS229 Machine Learning, Stanford CS230 Deep Learning, MIT RES.6-012 Introduction to Probability, Stanford EE178 Probabilistic Systems Analysis, Imperial College Mathematics for Machine Learning, IBM Applied Data Science Specialization, Statistical Rethinking 2026 (R. McElreath, Max Planck).