Research Profile
DTALab investigates trustworthy data-centric AI systems at the intersection of data management, natural language processing, large language models, machine learning and explainable AI.
Our research focuses on methods that connect data, models, evidence and explanations, with particular attention to systems that can be inspected, challenged and evaluated under realistic conditions. We work across textual, tabular and time-series data, developing methods, benchmarks and tools for data integration, evidence-grounded reasoning, robustness, responsible AI and decision support.

EVIDENCE · DOCUMENTS · REASONING
01 · Evidence-grounded NLP and LLMs
We develop NLP and LLM-based systems whose outputs can be traced back to explicit and verifiable evidence. Our research covers fact-checking and claim verification, evidence retrieval, source attribution, retrieval-augmented generation, and reasoning over complex textual and tabular information.
A particular focus is placed on table and document intelligence, including numerical and multi-table reasoning over corporate and sustainability reports. We study both how to retrieve the evidence required to answer a question and how to make the resulting reasoning process transparent and verifiable.

RECORDS · SCHEMAS · TABLES
02 · Data Integration and Entity Matching
We study methods for integrating heterogeneous data at the record, schema and table levels. A central research direction is Entity Matching, where the goal is to identify records referring to the same real-world entity across noisy and heterogeneous data sources.
Our work investigates automatic configuration of matching pipelines, transformer-based representations, model behaviour and sensitivity to data characteristics, as well as schema matching and large-scale table integration. The broader objective is to reduce the amount of manual configuration required to transform fragmented data sources into reliable integrated information.

SIGNALS · PATTERNS · ANOMALIES
03 · Interpretable Time-series Analysis
We develop methods for understanding complex multivariate time series, with particular attention to representations that expose the patterns driving model behaviour. Our research includes clustering, forecasting, anomaly and novelty detection, irregularly sampled data and self-supervised representations.
Applications range from industrial system monitoring to environmental and geoscientific data, where interpretability is particularly important for connecting detected patterns and anomalies to meaningful domain phenomena.

PERTURBATIONS · STRESS TESTS · RELIABILITY
04 · Semantic Robustness and Red-teaming
We investigate whether NLP and LLM systems remain reliable when their inputs, evidence or interaction context are deliberately modified. This includes adversarial triggers, semantic perturbations, misleading evidence, prompt-level variations and other forms of stress testing.
The objective is to determine whether model decisions are supported by stable semantic evidence or instead depend on lexical shortcuts, prompt artifacts or dataset-specific regularities. This research spans language models, fact-checking systems, retrieval-augmented pipelines and other evidence-driven AI systems.

EXPLANATIONS · FAIRNESS · RECOURSE
05 · Fairness, Auditability and Responsible AI
We develop methods that make AI decisions more understandable, auditable and actionable. Our work on explainability investigates both post-hoc and intrinsically interpretable approaches, identifying the attributes, terms and interactions that drive model predictions and studying the trade-off between explanation fidelity and human interpretability.
This perspective extends beyond Entity Matching to high-impact decision-support settings. In credit-risk modelling, we investigate counterfactual explanations that provide realistic and actionable alternatives to model decisions. Complementary research addresses fairness-aware learning, scalable mitigation strategies and auditing procedures for identifying undesirable model behaviour and analysing trade-offs between accuracy, fairness and operational constraints.

PREDICTION · PORTFOLIOS · DECISIONS
06 · Financial AI and Decision Support
We study machine-learning methods for financial prediction together with the evaluation procedures required to understand whether predictive improvements translate into better downstream decisions.
Our research connects heterogeneous financial data, forecasting and ranking models with portfolio construction and backtesting. A central objective is to develop reproducible evaluation frameworks that make models comparable across predictive tasks while explicitly accounting for their impact on portfolio-level outcomes.

CONTROL · INTERACTION · EXPERTISE
07 · Controllable LLM Systems for Education and Expert Workflows
We design LLM-based systems whose behaviour is constrained by explicit workflows rather than left entirely to unconstrained generation. Expert-defined phases, interaction strategies, assessment criteria and transition rules provide structured control over how the system interacts, evaluates information and decides what to do next.
This research is particularly relevant to education and other expert-guided settings, where users need to retain control over objectives and interaction policies. Applications include guided reading and intelligent tutoring, training activities and structured decision-support workflows.