About
I am a final-year Robotics and Mechatronics Engineering student at the University of Dhaka and an AI researcher working on multimodal intelligence, autonomous agents, and reliable decision-making. I am currently a Research Assistant at Cortex AI Lab and a Research Intern at the Data and Design Lab (CARS), University of Dhaka.
My research spans three closely connected directions. First, I study multimodal reasoning and vision-language models, with an emphasis on understanding how models perceive, reason over, and interact with complex visual information. Second, I work on agentic AI and autonomous decision-making, including LLM-based agents, multi-agent systems, investigation and information-seeking behavior, and the safety and robustness of agents operating in open environments. Third, I explore reinforcement learning and trustworthy AI, focusing on how representations, feedback, and evaluation shape the behavior and reliability of learning-based systems.
A recurring theme across my work is understanding when intelligent systems make good decisions, why they fail, and how we can evaluate those failures systematically. I am particularly interested in moving beyond benchmark accuracy toward behavioral evaluation, studying how models reason, adapt, investigate, cooperate, and behave under uncertainty, adversarial pressure, or limited information.
My recent research includes a NeurIPS 2026 main-track paper and two NeurIPS 2026 workshop papers, as well as publications in Findings of ACL 2026, Scientific Data, and the Journal of Hydrology: Regional Studies. Alongside fundamental AI research, I have worked on applied machine-learning systems spanning agriculture, power and critical infrastructure, and robotics.
Going forward, I am interested in developing AI systems that are not only more capable, but also more reliable, interpretable, and effective at reasoning and acting in complex environments.
News and Updates
-
Two papers accepted to NeurIPS 2026 workshops: The Surface You Test Is Not the Surface That Breaks at Agents in the Wild: Safety, Security, and Beyond (poster), and Unlearnable, or Unmeasured? at Transitioning from Pre-Training to Post-Training.
-
New preprint: Do Better Goal Representations Improve Goal-Conditioned Reinforcement Learning? is on arXiv.
-
🎉 Our paper MemeEconomy: Do LLM Agents Trade Ethics for Survival? is accepted to the NeurIPS 2026 main track.
-
PlantExpertVQA is published in Scientific Data (Nature Portfolio).
-
Our groundwater-recharge prediction paper is accepted at the Journal of Hydrology: Regional Studies.
-
Two new preprints on arXiv: The Surface You Test Is Not the Surface That Breaks and PhyDrawGen.
-
Thinking Like a Botanist is accepted to Findings of ACL 2026, with me as first author.
-
New preprint: Do Web Agents Investigate Before They Decide? is on arXiv.
-
Joined Cortex AI Lab, University of Dhaka, as a Research Assistant.
Selected Publications
All publications-
NeurIPS 2026
MemeEconomy: Do LLM Agents Trade Ethics for Survival?
A multimodal multi-agent market simulation across 300 events that tests whether LLM agents trade ethics for survival, with a compact verifier for auditing harmful agent behavior under pressure.
-
Findings of ACL 2026
Thinking Like a Botanist: Challenging Multimodal Language Models with Intent-Driven Chain-of-Inquiry
Introduces PlantInquiryVQA with 24,964 expert-curated images and 138,078 QA pairs. Structured chain-of-inquiry improves diagnostic correctness and reduces hallucination in vision-language models.
-
NeurIPS 2026 Workshop
Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR
Shows that “unlearnable” prompts in RL with verifiable rewards do improve, at about one third of the learnable rate, and that the difficulty labels defining them are far less reproducible than assumed.
-
NeurIPS 2026 Workshop
The Surface You Test Is Not the Surface That Breaks
Shows that prompt-injection vulnerability is a model-by-surface interaction. A surface-adaptive attacker improves attack success across 13 language models.
-
Preprint
Do Better Goal Representations Improve Goal-Conditioned Reinforcement Learning?
Finds that goal-representation quality barely affects offline goal-conditioned RL on OGBench navigation, while the state pathway is the bottleneck. Random Fourier positional encodings of the state more than double success on the hardest tasks.
-
Preprint
Do Web Agents Investigate Before They Decide?
A benchmark of 750 multi-step tasks isolating investigative competence across eight LLM agents. Reveals navigation-discovery gaps, contradiction failures, and investigative hallucination.
Research Experience
Full experience- Research Assistant, Cortex AI Lab, University of DhakaMar 2025 – Present
- Research Intern, Data and Design Lab (CARS), University of DhakaNov 2024 – Present
- Trainer, RobodemyMar 2025 – Dec 2025
- R&D Engineer, Tech Topia, DhakaNov 2023 – Jul 2025
Technical Skills
- Programming
- Python, C++, SQL, JavaScript, TypeScript
- Agentic AI
- LLM agents, tool use, multi-agent simulation, prompt-injection robustness, benchmark design
- Robot learning
- PyTorch, Gymnasium, ManiSkill / MS-HAB, reinforcement learning, policy evaluation
- Simulation & robotics
- Isaac Sim, Isaac Lab, MuJoCo, ROS 2, Gazebo, URDF, Xacro, kinematics, motion planning
- Perception & ML
- OpenCV, Transformers, vision-language models, TensorFlow, scikit-learn
- Hardware & tooling
- Arduino, ESP32, Autodesk Fusion 360, Docker, Git, Linux, MLflow
Honors and Awards
- Global Nominee, NASA Space Apps Challenge2024
- Runner-up, DU AI Challenge2025
- Runner-up, KUET Datathon2025
- Runner-up, Technocrats V2 IUBAT Hackathon2024
- Regional Champion, National High School Programming Contest2019
- Kaggle Expert, Multiple podium finishes in ML competitionsOngoing
I am always happy to talk about research and collaborations. LinkedIn is the best way to reach me.





