<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>PhD | Department of Knowledge Technologies</title>
	<atom:link href="https://kt.ijs.si/tag/phd/feed/" rel="self" type="application/rss+xml" />
	<link>https://kt.ijs.si</link>
	<description></description>
	<lastBuildDate>Fri, 19 Jun 2026 08:20:34 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.0</generator>

<image>
	<url>https://kt.ijs.si/wp-content/uploads/2021/04/cropped-Knowledge-tehnologies_favicon_s-32x32.png</url>
	<title>PhD | Department of Knowledge Technologies</title>
	<link>https://kt.ijs.si</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Taja Kuzman Pungeršek successfully defended her doctorate thesis</title>
		<link>https://kt.ijs.si/news/taja-kuzman-pungersek-successfully-defended-her-doctorate-thesis/</link>
		
		<dc:creator><![CDATA[Anja Glusic]]></dc:creator>
		<pubDate>Fri, 15 May 2026 11:34:48 +0000</pubDate>
				<category><![CDATA[News]]></category>
		<category><![CDATA[PhD]]></category>
		<guid isPermaLink="false">https://kt.ijs.si/?p=7438</guid>

					<description><![CDATA[Taja Kuzman Pungeršek successfully defended her doctorate thesis titled Robust Multilingual Automatic Genre Identification in Texts. Congratulations! Abstract: Collecting texts from the web has significantly accelerated the creation of large text datasets which are essential for the development of advanced language technologies, including large language models. However, because these texts are gathered automatically, their linguistic [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Taja Kuzman Pungeršek successfully defended her doctorate thesis titled<em> Robust Multilingual Automatic Genre Identification in Texts.</em></p>
<p>Congratulations!</p>
<p>Abstract:</p>
<p>Collecting texts from the web has significantly accelerated the creation of large text datasets<br />
which are essential for the development of advanced language technologies, including large<br />
language models. However, because these texts are gathered automatically, their linguistic<br />
and functional characteristics are largely unknown, limiting their reliable use in research<br />
and applications. Automatic genre identification, a text classification task that categorizes<br />
texts into specific genre categories, provides key insights into such large-scale text<br />
collections and enables their filtering for applications in language technology and linguistic<br />
research.<br />
This thesis advances automatic genre identification for web-scale multilingual text data<br />
through the creation of novel genre schemata, the development of manually-annotated<br />
genre datasets, and the exploration of robust machine learning methods. First, the<br />
study introduces new genre schemata designed to improve annotation reliability and<br />
cross-schema comparability, thereby addressing a key limitation of prior work where<br />
incompatible schemata hindered meaningful comparison across studies. We demonstrate<br />
that high-quality manual annotation with acceptable inter-annotator agreement can<br />
be achieved through a carefully designed genre schema, detailed guidelines, and the<br />
employment of expert annotators.<br />
To address the lack of high-coverage genre datasets in our target languages, we develop<br />
high-quality, manually-annotated genre datasets in Slovenian and English, as well as<br />
multilingual test collections spanning eleven typologically diverse languages and different<br />
scripts. Building on the test datasets, we establish the AGILE benchmark, which enables<br />
standardized, reproducible cross-dataset and cross-lingual evaluation of genre classifiers.<br />
The benchmark also supports a systematic comparison of modern large language models<br />
across languages.<br />
Using this infrastructure, we develop genre classifiers that achieve robust performance<br />
across various languages, datasets, and evaluation settings. To ascertain which machine<br />
learning methodology offers the most robust generalization, we experiment with a range<br />
of text classification techniques, ranging from traditional non-neural machine learning<br />
methodologies to cutting-edge approaches based on large language models. Our findings<br />
indicate that a BERT-like model, fine-tuned on our newly developed training dataset,<br />
achieves state-of-the-art performance across various datasets and languages, including languages<br />
using non-Latin scripts and languages not closely related to the fine-tuning languages.<br />
The resulting model is publicly released, providing a practical tool for large-scale<br />
automatic genre annotation of multilingual web text collections.<br />
Furthermore, we propose the LLM Teacher-Student Framework, a novel approach for<br />
training text classifiers without manually-annotated data. By leveraging a large language<br />
model to generate training labels, this method enables scalable, cost-efficient development<br />
of genre classifiers, particularly for low-resource languages and specialized domains, and is<br />
applicable beyond genre identification to other text classification tasks.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
