Research Topic:
Robust Reinforcement Learning for Large Language Models: Geometry-Regulated Policy Optimization, Process Reward Integration, and Agentic AI
Supervisors:
Prof. Sahar Vahdati (LUH), Prof. Jens Lehmann (TUD)
Thesis Abstract:
My research develops robust reinforcement learning methods for large language models by improving training stability, reasoning accuracy, and efficiency. It introduces geometry-regulated policy optimization and verifiable process rewards to reduce reward hacking while enhancing step-by-step reasoning. The project also extends reinforcement learning to tool-using AI agents and investigates efficient inference methods for faster and more cost-effective language model deployment.
Publications:
•Mohammad Rezaei, Jens Lehmann, and Sahar Vahdati. “LLM Reasoning with Process Rewards for Outcome-Guided Steps”. In: Proceedings of the International Joint Conference on Neural Networks (IJCNN). Presented at IJCNN 2026. 2026. url: arxiv.org/abs/2604.02341.
Conference Presentations:
•International Joint Conference on Neural Networks (IJCNN 2026): LLM Reasoning with Process Rewards for Outcome-Guided Steps.
