Hi there! I am Tianyu Liu (刘天宇), a fourth-year joint Ph.D. student at USTC and Shanghai AI Laboratory, supervised by Xiao Sun.

My research focuses on efficient inference for LLMs, especially speculative decoding. I proposed PEARL (ICLR 2025), the first parallel speculative decoding framework, and D-cut, which is now shipped in Tencent's open-source AngelSlim and followed by DeepSeek's DSpark. I am always open to collaborations on related inference topics.

I expect to graduate in 2027 and am currently seeking job opportunities in efficient LLM inference and AI systems. Please feel free to contact me at tianyu_liu@mail.ustc.edu.cn.

🔥 News

  • 2026.08KVShot is accepted to the COLM 2026 Workshop on Efficient Reasoning!
  • 2026.07Released AngelSpec, the Tencent Hunyuan technical report on real-world high performance inference with speculative decoding.
  • 2026.07D-cut is on arXiv, reaching up to 3.0× speedup over autoregressive decoding on MoE models under high concurrency.
  • 2026.06Introduce D-cut, an adaptive verification-depth pruning method that accelerates speculative decoding at high concurrency.
  • 2026.05Released two new preprints on speculative decoding: KVShot and Graft!
  • 2026.04Double is accepted to ACL 2026 main conference as an oral presentation, and LogitSpec is accepted to ACL 2026 Findings!
  • 2026.01SpecBranch is accepted to ICLR 2026!
  • 2026.01Released four preprints: TALON, KALE, Double, and HIPPO!
  • 2025.10Released nano-PEARL, an engineering follow-up to PEARL with multi-GPU draft-target disaggregation.
  • 2025.07Preprint LogitSpec to arXiv, a training-free retrieval-based speculative decoding method!
  • 2025.01PEARL is accepted to ICLR 2025!
  • 2025.01One paper accepted to NAACL 2025. Thanks for the carry of Qitan!
  • 2023.09REST is accepted to NeurIPS 2023!

📝 Selected Publications

: corresponding author; *: equal contribution

ICLR 2025
PEARL illustration

Featured PEARL: Parallel Speculative Decoding with Adaptive Draft Length

Tianyu Liu, Yun Li, Qitan Lv, Kai Liu, Jianchen Zhu, Winston Hu, Xiao Sun

[Paper]  [Project]  [Code]


🚀 Introduce nano-PEARL: a Draft-Target Disaggregation LLM Serving System via Parallel Speculative Decoding!

Preprint
D-cut illustration

Featured D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding

Tianyu Liu, Yuhao Shen, Rui Cen, Junhan Shi, Jiebin Zhang, Guangshuo Qin, Hong Liu, Song Liu, Guanghua Yu, Jianchen Zhu

[Paper]  [Code]  [Docs]


🚀 Followed by DeepSeek DSpark!

COLM 2026 Workshop
KVShot illustration

When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding?

Tianyu Liu, Yuhao Shen, Xinyi Hu, Baolin Zhang, Hengxin Zhang, Jun Dai, Jun Zhang, Shuang Ge, Lei Chen, Yue Li, Mingcheng Wan

[Paper]

Preprint
TALON illustration

TALON: Confidence-Aware Speculative Decoding with Adaptive Token Trees

Tianyu Liu*, Qitan Lv*, Yuhao Shen, Xiao Sun, Xiaoyan Sun

[Paper]

ACL 2026 Findings
LogitSpec illustration

LogitSpec: Accelerating Retrieval-based Speculative Decoding via Next Next Token Speculation

Tianyu Liu*, Qitan Lv*, Hao Li, Xing Gao, Xiao Sun, Xiaoyan Sun

[Paper]  [Code]

NeurIPS 2023
REST illustration

Learning Rule-Induced Subgraph Representations for Inductive Relation Prediction

Tianyu Liu, Qitan Lv, Jie Wang, Shuling Yang, Hanzhu Chen

[Paper]  [Code]

🗂 Other Publications

Tech Report AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding

Hong Liu, Rui Cen, Junhan Shi, Guangshuo Qin, Jiebin Zhang, Tianyu Liu, Runzhi Fan, Guoliang Zhao, Ruobing Xie, Kai Zhang, Song Liu, Guanghua Yu, Jianchen Zhu

[Paper]  [Code]  [Blog]

arXiv 2026 Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding

Yuhao Shen, Tianyu Liu, Xinyi Hu, Quan Kong, Baolin Zhang, Jun Dai, Jun Zhang, Shuang Ge, Lei Chen, Yue Li, Mingcheng Wan, Cong Wang

[Paper]

arXiv 2026 HIPPO: Accelerating Video Large Language Models Inference via Holistic-aware Parallel Speculative Decoding

Qitan Lv*, Tianyu Liu*, Wen Wu, Xuenan Xu, Bowen Zhou, Feng Wu, Chao Zhang

[Paper]

arXiv 2026 KALE: Enhancing Knowledge Manipulation in Large Language Models via Knowledge-aware Learning

Qitan Lv*, Tianyu Liu*, Qiaosheng Zhang, Xingcheng Xu, Chaochao Lu

[Paper]

ACL 2026 Main (Oral) Double: Breaking the Acceleration Limit via Double Retrieval Speculative Parallelism

Yuhao Shen, Tianyu Liu, Junyi Shen, Jinyang Wu, Quan Kong, Li Huan, Cong Wang

[Paper]  [Code]

ICLR 2026 Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism

Yuhao Shen, Junyi Shen, Quan Kong, Tianyu Liu, Yao Lu, Cong Wang

[Paper]  [Code]

NAACL 2025 Exploiting Edited Large Language Models as General Scientific Optimizers

Qitan Lv*, Tianyu Liu*, Hong Wang

[Paper]

📖 Education

  • 2022.09 - Present, Ph.D. in Information and Communication Engineering, University of Science and Technology of China.
  • 2018.09 - 2022.06, B.Eng. in Computer Science and Technology, Central University of Finance and Economics.

🖥️ Academic Service

  • Conference reviewer for ICLR'25, ICLR'26, WWW'25, NeurIPS'25, NeurIPS'26.