Hi, I’m Hassan.

I’m a PhD student in Computer Science at the Ubiquitous Knowledge Processing (UKP) Lab, TU Darmstadt, supervised by Prof. Iryna Gurevych.

I work on NLP and LLMs, with a focus on reliable knowledge access and LLM safety. I’m interested in how systems retrieve and use evidence—and what happens when important evidence is missing.

My recent work includes CORE-T, a framework for retrieving coherent sets of tables for text-to-SQL. I’m also investigating the robustness of retrieval-augmented and agentic systems to incomplete evidence.

Previously, I was an AI researcher in DFKI’s Educational Technology Lab, supervised by Dr. Miloš Kravčík. I worked on educational chatbots, retrieval-augmented mentoring, and adaptive dialogue.

I completed my MSc in Computer Science at Saarland University. My master’s thesis on cross-domain neural entity linking was supervised by Prof. Dietrich Klakow at Saarland University, in collaboration with Bosch Center for AI (BCAI), where I worked with Dr. Heike Adel, Dr. Mohamed H. Gad-Elrab, Dragan Milchevski, and Dr. Jannik Strötgen.

News

CORE-T accepted to EMNLP 2026 (Main).

Released the CORE-T preprint on coherent multi-table retrieval for text-to-SQL.

Started tutoring Introduction to LLMs at TU Darmstadt for winter semester 2025–26.

Retrieval-Augmented Chatbots for Scalable Educational Support in Higher Education published in the GenAI-LA @ LAK 2025 proceedings.

Presented our educational chatbot work at GenAI-LA @ LAK 2025.

Started my PhD at UKP Lab, TU Darmstadt.

Scalable Mentoring Support with a Large Language Model Chatbot published at ECTEL; received the Second Best Demo Paper award.

Our work on LLM-based learning support in educational sciences appeared at DELFI.

Our adaptive dialogue management paper published in ACM UMAP Adjunct ’24.

Research

Current directions

Evidence-aware, safer language models

My ongoing SafeLLMs research studies strategic missingness: how omitted evidence can affect the conclusions of LLM-based systems. I’m exploring evidence coverage, retrieval, and evaluation in retrieval-augmented generation and agentic fact-checking.

Ongoing research

Retrieval over structured knowledge

How can a system find not just individually relevant tables, but a set that works together? In CORE-T, we study training-free, join-coherent table retrieval for open-book text-to-SQL over large, heterogeneous collections.

CORE-T publication ↓

Adaptive dialogue & learning support

My earlier work explored LLM-based mentoring and dialogue systems that adapt to their users, alongside entity linking across domains.

Publications

Selected work

Background

Research & industry

Jan 2025 – present

Doctoral Researcher · TU Darmstadt

UKP Lab · Supervised by Iryna Gurevych

  • Developed CORE-T for training-free, coherent table retrieval, evaluated across BIRD, Spider, MMQA, and BEAVER; accepted to EMNLP 2026 Main.
  • Investigating evidence coverage and strategic missingness in retrieval-augmented and agentic systems, alongside tutoring and master’s thesis supervision.
Jan 2023 – Dec 2024

AI Researcher · DFKI

Educational Technology Lab

  • Led two projects and supervised two students; developed a graduate-course chatbot answering student queries with 87% accuracy.
  • Combined hybrid retrieval and reranking with LangGraph mentoring workflows and small open-source models on Azure; benchmarked dialogue adaptation to emotional state and demographics.
May 2022 – Aug 2022

Applied Scientist Intern · Bosch Center for AI

NLP & Semantic Reasoning

  • Transferred neural entity-linking research to industrial data, achieving 77% end-to-end top-3 recall on a large domain-specific dataset.
  • Fine-tuned models on the in-house GPU cluster and refactored, tested, and documented production-level ML code.
Jun 2021 – Jan 2022

Master’s Thesis Student · Bosch Center for AI

In collaboration with Saarland University

  • Used context-aware BERT embeddings in a joint vector space to link entities across Wikipedia and domain-specific knowledge bases.
  • Improved top-1 average precision by 9% and top-10 MAP by 20% across four domain-specific knowledge bases; the work led to a RepL4NLP @ ACL 2022 publication.
Nov 2020 – May 2021

Research Assistant · Max Planck Institute for Informatics

Database & Information Systems

  • Developed entity set-expansion prototypes using Wikipedia lists to identify diverse peer groups for entities.
  • Achieved a 3× faster runtime through efficient sparse matrix multiplication.
Aug 2019 – Feb 2020

Software Development Engineer Intern · Amazon

Fulfillment Acceleration · Luxembourg

  • Maintained an AWS-based web simulation tool as a full-stack engineer to model delivery speed for Prime customers.
  • Supported fulfillment analysis and reporting, and maintained server infrastructure and team tools in an Agile environment.

Education

Jan 2025 – present

PhD Computer Science

TU Darmstadt · UKP Lab

Supervised by Prof. Iryna Gurevych · NLP, retrieval, and LLMs

Oct 2018 – Sep 2022

MSc Computer Science

Saarland University · Overall GPA: 1.40 / 1.00

Thesis: Cross-Domain Neural Entity Linking

Supervisor: Prof. Dietrich Klakow

Sep 2013 – Sep 2018

BSc Computer & Communication Engineering

Alexandria University · Overall GPA: 3.96 / 4.00 · First Class Honours

Thesis: Car License Plate Arabic Numbers & Letters Recognition

Supervisor: Dr. Marwan Torki

Awards

Second Best Demo Paper · ECTEL 2024

For Scalable Mentoring Support with a Large Language Model Chatbot.

First Class Honour Degree · Alexandria University

Recognized for outstanding academic performance in Computer & Communication Engineering.

Projects

Implementations & experiments

SmolLM + reinforcement learning ↗

Implemented a 135M-parameter language model and fine-tuned it for grammatical error correction.

  • Built the architecture with rotary positional embeddings, KV caching, grouped-query attention, RMSNorm, and SwiGLU.
  • Fine-tuned on Grammarly CoEdIT and applied RLAIF through DPO; provided a Colab notebook covering implementation, training, and evaluation.

PyTorch · Hugging Face · TRL

LinguaLexMatch ↗

Compared multilingual embeddings, TF-IDF with Multinomial Naive Bayes, and fine-tuned transformers for language identification across 20 languages.

Multilingual embeddings · Scikit-learn · Transformers

Teaching & Service

Tutoring · Introduction to LLMs

TU Darmstadt · NLP foundations, LLM architectures and training, prompting, fine-tuning, retrieval-augmented generation, and hands-on model evaluation.

  • Winter semester 2025–26: tutoring through practical exercises, programming assignments, and student support.
  • Winter semester 2026–27: upcoming tutoring, including preparation of practical sessions and learning materials.

Master’s Thesis Supervision

  • Abdelrahman Abdelgawad — extending CORE-T to multimodal retrieval: finding coherent evidence across modalities for downstream question answering.

Academic Reviewing

Get in touch

For research questions or to discuss related work, you can reach me at .