Masterarbeit
Title: Enhancing Scholarly Knowledge Graphs: Ontology-Guided Extraction of Research Results from Scientific Literature
Research Area
Web Engineering
Students
Advisers
Description
In the rapidly expanding field of computer science, the sheer volume of scientific literature published daily presents a significant bottleneck for researchers attempting to identify relevant methodologies, datasets, objectives, and experimental results. This challenge is heavily compounded by the unstructured nature of academic publications, which contain complex layouts, nested text elements, mathematical expressions, tables, and embedded programming code. Traditional text extraction methods struggle to accurately segment these heterogeneous formats, and rule-based systems suffer from poor generalizability. Furthermore, semantic inconsistencies where distinct expressions such as “SVM” and “Support Vector Machine” are used to describe the same concept create fragmented data landscapes, hindering automated statistical trends and aggregate analysis. Consequently, researchers face immense difficulties in rapidly grasping advancements, discovering trends, and making meaningful comparisons between scholarly articles.
To address these challenges, this thesis proposes an automated semantic framework that combines advanced document parsing, state-of-the-art Large Language Models (LLMs), and ontology-guided Knowledge Graph construction to extract and organize research results from the scientific literature. Input documents are systematically collected and parsed using Python libraries such as PyMuPDF, Requests, and BeautifulSoup to separate textual content, code blocks, and data tables. A comparative evaluation of frontier LLMs will be conducted using advanced prompt engineering to identify and extract research result entities and their relationships. To ensure semantic consistency and resolve synonymous terms, the extracted research results are aligned with established domain ontologies, including the Open Research Knowledge Graph, Computer Science Ontology, Semantic Survey Ontology, and Software Ontology. The aligned data is then transformed into RDF triples and integrated into a unified, searchable, and semantically enriched Knowledge Graph.
The objective of this master thesis is to develop, implement, and evaluate this automated semantic framework to support the efficient comparison and discovery of research results. Specifically, the thesis aims to: Design and implement a robust ingestion and extraction pipeline capable of parsing complex, multi-layered document structures and isolating natural language from code and mathematical expressions. Develop a semantic harmonization module that resolves synonymy and aligns extracted entities with the ORKG, CSO, SemSur and SWO ontologies. Construct and validate a unified Knowledge Graph to allow semantic search and structured querying. The proposed scholarly Knowledge Graph construction framework will be evaluated with respect to quantitative extraction performance and qualitative utility. The quantitative evaluation will assess the extracted RDF triples against a manual gold-standard evaluation and compare the results with an existing approach, such as ORKG, on the same input. Finally, a human error analysis will categorize failure modes (e.g., hallucinations, boundary mismatches) alongside SPARQL-based competency questions to verify downstream graph utility.