Text and data mining refers to methods for the automated analysis of large volumes of text and data, with the aim of identifying hidden patterns and correlations, as well as systematically extracting information that is obvious to humans and making it available in a structured format. This field is particularly relevant to information and library science, as it opens up new possibilities for efficiently cataloguing, analysing and making large digital collections available in a way that is tailored to specific target groups.
Methods drawn primarily from statistics, natural language processing, machine learning and artificial intelligence (AI) are used to analyse texts and large volumes of data. At the same time, natural language processing and machine learning form the basis of large language models, which constitute the backbone of modern AI applications. This content is taught systematically across several modules and sub-modules and is firmly embedded in the programme’s curriculum.
In the ‘Fundamentals of AI’ module (BIM-137) in the third semester, you will learn, on the one hand, what possibilities AI systems offer, how you can utilise them, and which prompting strategies you can employ for this purpose. In the sub-module ‘Fundamentals of Computational Linguistics’, you will learn how languages are structured, how information can be extracted from texts, and how language models are trained and can generate text.
In the fifth semester, in the ‘Information Retrieval’ module (BIM 256), you will learn how modern search engines work and gain a detailed understanding of how texts are analysed and indexed, and how search results are effectively ranked. In doing so, you will work with both traditional methods and innovative approaches based on large language models. There is a particular focus on practical application: you will develop a small search engine or a RAG system yourself, thereby gaining direct insights into current technologies. RAG stands for Retrieval Augmented Generation and refers to an approach in which generative language models are combined with external information sources to generate well-founded answers based on relevant documents.
In the 6th semester, you will deepen your knowledge in the Text and Data Mining module (BIM 267), which is divided into the sub-areas of text mining and data mining. The focus is on fundamental machine learning techniques as well as key methods for the automated processing and analysis of text and structured data. Among other things, you will learn how to develop and fine-tune models for classification tasks, and how modern language models can be fine-tuned and deployed for specific applications in text and data mining.