Mohammad Rezaei

 

Center: NHR@TUD

 

Research Topic: Robust Reinforcement Learning for Large Language Models: Geometry-Regulated Policy Optimization, Process Reward Integration, and Agentic AI

Research Topic:

Robust Reinforcement Learning for Large Language Models: Geometry-Regulated Policy Optimization, Process Reward Integration, and Agentic AI

 

 

Supervisors:

Prof. Sahar Vahdati (LUH), Prof. Jens Lehmann (TUD)

 

Thesis Abstract:

My research develops robust reinforcement learning methods for large language models by improving training stability, reasoning accuracy, and efficiency. It introduces geometry-regulated policy optimization and verifiable process rewards to reduce reward hacking while enhancing step-by-step reasoning. The project also extends reinforcement learning to tool-using AI agents and investigates efficient inference methods for faster and more cost-effective language model deployment.

 

Publications:

•Mohammad Rezaei, Jens Lehmann, and Sahar Vahdati. “LLM Reasoning with Process Rewards for Outcome-Guided Steps”. In: Proceedings of the International Joint Conference on Neural Networks (IJCNN). Presented at IJCNN 2026. 2026. url: arxiv.org/abs/2604.02341.

 

 

Conference Presentations:

•International Joint Conference on Neural Networks (IJCNN 2026): LLM Reasoning with Process Rewards for Outcome-Guided Steps.