Tianshi Zheng 郑天石 Stone
Bibliography
Hello there!
I’m a Ph.D. candidate in Computer Science at the Hong Kong University of Science and Technology (HKUST) starting from 2024, advised by Prof. Yangqiu Song. Prior to that, I obtained two bachelor’s degrees in Computer Science and General Business Management from the Dual Degree Program at the same institution with the highest distinction.
News
- Sep 2026: Our latest work on scientific ideation: AgentIdeaBench has been published. with benchmarking code AgentIdeaBench and LeaderBoard!
- Aug 2026: Six research papers accepted to EMNLP 2026!
- May 2026: Our latest work on frontier scientific reasoning: SciResearcher has been published.
- Apr 2026: Four research papers accepted to ACL 2026 Main Conference!
- Jan 2026: Two research papers accepted to ICLR 2026!
- Nov 2025: One research paper accepted to TMLR!
- Oct 2025: Successfully passed my PhD Qualification Defense! Thanks to my committee members Junxian and May!
- Oct 2025: Our latest work on scientific discovery: NewtonBench has been published, with benchmarking code released in github: NewtonBench .
- Aug 2025: Two research papers accepted to EMNLP 2025 Main Conference!
- May 2025: Our survey paper on Large Language Models in Scientific Discovery has been published, with an awesome github repository: Awesome-LLM-Scientific-Discovery . We welcome community contributions through pull requests!
- May 2025: Four research papers accepted to ACL 2025!
Research Interests
I work on Machine Learning and Natural Language Processing, with a focus on the full lifecycle of AI-Driven Autonomous Scientific Discovery and Research.
- Theory, Taxonomy and Survey: LLM-SD Survey, Discovery Geometry.
- Ideation and Hypothesis Generation: AgentIdeaBench.
- Hypothesis over Structural Data: AbductiveKGR, CtrlHGen, HypoAgent.
- Hypothesis as Inductive/Abductive Discovery: LogiDynamics, The Curse of CoT, Robust Rule Induction, Legal Rule Induction.
- Scientific Deep Research and Reasoning: SciResearcher, CK-Pro.
- Scientific Principle Discovery: NewtonBench.
- AutoResearch, AI4AI, AI4Math and AI4Science.
Before stepping into Agent4Research, I also spent time in logical reasoning, graph learning, and games.
Technical-wise, I focus on agentic post-training (SFT/RL) and environment scaling to improve scientific research abilities of foundation models.
Selected Papers
SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning.
Tianshi Zheng, Rui Wang, Xiyun Li, Yangqiu Song, Tianqing Fang.
Under Review [pdf]NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents.
Tianshi Zheng*, Kelvin Tam*, Newt Nguyen*, Baixuan Xu, Zhaowei Wang, Jiayang Cheng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Tianqing Fang, Yangqiu Song, Ginny Y Wong, Simon See.
ICLR 2026 [pdf], [code]The Curse of CoT: On the Limitations of Chain-of-Thought in In-Context Learning.
Tianshi Zheng*, Yixiang Chen*, Chengxi Li*, Chunyang Li, Qing Zong, Haochen Shi, Baixuan Xu, Yangqiu Song, Ginny Y Wong, Simon See.
TMLR 2025 [pdf], [code]From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery.
Tianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Zihao Wang, Yangqiu Song.
EMNLP 2025 Main Conference [pdf], [code]LogiDynamics: Unraveling the Dynamics of Logical Inference in Large Language Model Reasoning.
Tianshi Zheng, Jiayang Cheng, Chunyang Li, Haochen Shi, Zihao Wang, Jiaxin Bai, Yangqiu Song, Ginny Y Wong, Simon See.
EMNLP 2025 Main Conference [pdf], [code]Enhancing Transformers for Generalizable First-Order Logical Entailment.
Tianshi Zheng*, Jiazheng Wang*, Zihao Wang, Jiaxin Bai, Hang Yin, Zheye Deng, Yangqiu Song, Jianxin Li.
ACL 2025 Main Conference [pdf], [code]KnowShiftQA: How Robust are RAG Systems when Textbook Knowledge Shifts in K-12 Education?
Tianshi Zheng*, Weihan Li*, Jiaxin Bai, Weiqi Wang, Yangqiu Song.
ACL 2025 Main Conference [pdf], [code]
Awards
- Hong Kong PhD Fellowship 2024
- HKUST RedBird PhD Scholarship 2024
- HKUST Academic Achievement Medal (~top 1% graduate)
- HKUST Dean's Lists (7 semesters)
Talks
- 2026 Aug: NVIDIA. Towards Generalizable and Self-Evolving Agents for Scientific Discovery.
- 2025 Dec: Shanghai AI Lab. NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents.
- 2025 Nov: SJTU & Suzhou Lab. From Automation to Autonomy: The Paradigm Shift of LLMs' Role in Scientific Discovery.
- 2025 Aug: NVIDIA. Thinking vs. Intuition: Between System I / System II Paradigms of LLM Reasoning.
Academic Service
Conference Reviewer: ACL Rolling Review 2023-2026, KDD 2024, NLPCC 2024, ICLR 2025, COLM 2025-2026.
