LLM post-training · alignment · evaluation · agents 🤖 — MSc @ BJTU
- 🎓 MSc in Software Engineering @ Beijing Jiaotong University; BEng @ Nantong University.
- 🔬 I work on LLM post-training, alignment, and evaluation, and build LLM agents and developer tools.
- 🛠️ I contribute fixes to open-source training, evaluation, and agent frameworks, with a focus on numerical stability, metric correctness, and reliable runtime behavior.
- 📝 Co-authored a research paper.
- ✍️ I write notes & blog posts at excelius.xyz.
- ByteDance — Multimodal LLM Algorithm Intern · present
- Baidu — LLM Post-training Algorithm Intern
- Tsinghua University, Institute of Vehicle Power & Intelligent Energy — LLM Application & Full-stack Intern
45 merged PRs across external open-source repositories, covering code fixes, tests, and documentation. Counts below include merged PRs only, as of September 13, 2026.
- OpenRLHF — 7 merged PRs: Improved PPO numerical stability, reward computation, checkpoint recovery, evaluation batch handling, and multi-turn rollout truncation. Selected PRs: k3 KL gradients #1335 · FP32 reward shaping #1334 · checkpoint recovery #1333 · rollout truncation #1327.
- ms-swift — 6 merged PRs: Fixed NLG metric aggregation, inference error propagation, prompt token accounting, SSE handling, and training integration compatibility. Selected PRs: empty prediction scoring #9962 · worker errors #10093 · prompt usage #10094 · Megatron integration #10025.
- EvalScope — 6 merged PRs: Corrected ASR word error rates, streaming latency measurements, multimodal image inputs, and Terminal-Bench reward validation; also updated evaluation configuration and documentation. Selected PRs: ASR scoring #1722 · TTFT / ITL #1645 · image inputs #1618 · trial rewards #1610.
- OpenAI Agents Python — 3 merged PRs: Fixed session deletion under cancellation, model-provider cleanup, and type annotations for variadic tool arguments. PRs: session cleanup #4790 · provider lifecycle #4785 · tool arguments #4655.
Other merged contributions include Atomic Agents (4), LiveKit Agents (2), MCP Servers (1), OpenAI Agents JS (1), and Axolotl (1). Examples: MCP resource templates · turn cancellation · UTF-8 file reads · hosted MCP outputs · activation checkpointing compatibility.
- Languages: Python, C++, TypeScript / JavaScript, Rust
- LLM / Post-training: PyTorch, Hugging Face Transformers, LLaMA-Factory, ms-swift, vLLM
- Alignment / Evaluation: SFT, LoRA, DPO, PPO, GRPO, reward modeling, LLM-as-a-Judge, EvalScope
- Agents: LangChain, LangGraph, tool calling, ReAct, plan-and-execute workflows
- Applications / Infra: React, Vue, FastAPI, Tauri, multi-node multi-GPU training, Git, Docker
- Repolane — A GitHub desktop workspace for browsing code, reviewing pull requests, and managing repository work, built with React, TypeScript, Rust, and Tauri. Repository:
harbor; actively developing, with no packaged public release yet. - Self-DeepResearch — An iterative research agent built with LangGraph and Tavily: plan → search → review → report, with a Vue frontend and FastAPI streaming backend.
- marginalia — A paper-reading skill for Claude Code / Codex that reconstructs a paper's reasoning, examines its assumptions, and publishes structured notes with figures and formulas to a Feishu knowledge base.
- dive-into-transformer-pytorch — A Transformer language model implemented in PyTorch and trained on Dream of the Red Chamber, with DDP multi-GPU training, checkpoint saving, and training-curve visualization.
- BERT_BiLSTM_CRF — Chinese named-entity recognition with BERT + BiLSTM + CRF, developed for BJTU NLP coursework.
|
|
|
||||
|
|
|
|
|||
|
|
|||||



