I build AI platforms and research infrastructure that integrate voice, gesture, language, and physiology, advancing the science of human signal at scale. I work in the open, so that what we learn about human beings remains a shared resource.
Multimodal Open-Science AI-Powered Integrative Computation
Humans don't communicate in channels. They speak, move, breathe, and feel simultaneously. MOSAIC is an AI-powered data science platform for integrating physiological, behavioral, linguistic, clinical, and subjective data into a unified research infrastructure. Built for science, designed for the full complexity of the human signal. It is open by design: its methods, standards, and tooling are released to the wider research community, because infrastructure for understanding people should serve the public, not be held privately.
See All Research ↓Guiding research, supervising junior researchers and teaching in the domain of multimodal language and behavior processing using video, speech and natural language processing (NLP).
Multimodal Open-Science AI-Powered Integrative Computation
AI-powered data science platform for processing human signal, involving physiological, behavioral, linguistic, clinical, and subjective data.
Language Data Across Disciplines
Development of teaching materials and workshops on data analytics and quantitative and computational skills for language data, run jointly across nine Swiss institutions. The full programme is documented on the CLARIN-CH project page.
Digital Audio-Visual Annotation Lab
Development of AI-powered software for automatic audio-visual and multimodal data analysis, aimed at academia and industry. The lab and its aims are described on the Digital Society Initiative pages.
Multimodal Audio-Visual Analysis Cluster
Cluster for infrastructure interoperability and data exchange, connecting VideoScope, TIB-AV-A and VIAN through shared formats and APIs so that audio-visual research data can move between them under FAIR principles. The resulting exchange format for video annotations is developed in the open as mava-exchange, and the cluster is described in the announcement at LiRI.
Computational Analysis of Personal Identity in Interaction: Recognition and Ethics
Focus on explainable machine learning for personal identity patterns in speech, facial expressions and gestures.
Application and research methodology development, implementing and training models for image and speech recognition, natural language processing (NLP). The platform is live at lcp.linguistik.uzh.ch — with CatchPhrase, SoundScript and VideoScope for text, audio and video corpora — and developed in the open on GitHub.
(Dis-)entangling Traditions on the Central Balkans: Performance and Perception
Creation of electronic language resources and corpus-based analysis of linguistic variation. The resulting Spoken Torlak Dialect Corpus — semi-orthographic transcripts of 86.5 hours of fieldwork recordings — is openly published through CLARIN.SI, where it can also be queried directly in the browser. Funded by ERA.Net RUS Plus under FP7/Horizon 2020.
Developing the corpus of the Torlak dialect, collected in the field across the Timok area and later released as the openly available Spoken Torlak Dialect Corpus, searchable via CLARIN.SI.
Corpus linguistics, digital humanities, fieldwork, dialectology, digital lexicography. Author of the Bunjevac dialect corpus.
Practice and teaching alongside research that turns the same multimodal methods — language, speech, physiology — toward first-person experience.
A human development initiative grounded in truth seeking and the reduction of suffering, teaching a variety of meditation tools in the spirit of pragmatic dharma. More at meditative.dev.
NLP analysis of personal experience reports and meditative phenomenology reports gathered during an advanced meditation retreat.
Participant in an advanced meditation retreat during which the Riddle Lab collected EEG and speech data.
Meditation guidance and multimodal analysis expertise for a robotic system delivering personalized meditation instruction through neurofeedback.
NLP, AI and statistical analysis for an exploratory project investigating the process of awakening through large text collections and social network data. Member of the Emergent Phenomenology Research Consortium, a multidisciplinary alliance studying emergent phenomena with scientific and clinical methods.
170 citations · h-index 7 · i10-index 5 (Google Scholar, August 2026)
Full profile on Google Scholar →
A Modular Approach Towards Representing Multi-Dimensional Corpora: Theoretical Considerations behind the LiRI Corpus Platform / Linguistic Data Environment
Submitted to Language Resources and Evaluation
Micro-Expression-Aware Avatar Fingerprinting via Inter-Frame Feature Differencing
Proceedings of FG2026, IEEE Biometrics
Spoken Corpora
International Encyclopedia of Language and Linguistics, 3rd edition. Elsevier
Towards Language-Independent Face-Voice Association with Multimodal Foundation Models
Odyssey 2026
TidyVoice 2026 Challenge Evaluation Plan
Interspeech 2026
The Swiss FAIR-compliant ecosystem of infrastructures 2.0
CLARIN Annual Conference Proceedings
Beyond Appearance: Transformer-based Person Identification from Conversational Dynamics 2nd Best Paper
15th International Conference on Computer and Knowledge Engineering (ICCKE 2025), IEEE
LiRI Corpus Platform: Demonstration of a Web-Based Infrastructure for Multimodal Corpus Analysis
Interspeech 2025
NLP for preserving Torlak, a vulnerable low-resource Slavic language
Proceedings of the 31st International Conference on Computational Linguistics. ACL Anthology
CL-UZH submission to the NIST SRE 2024 Speaker Recognition Evaluation
NIST SRE 2024
Team Switzerland Submission to NIST SRE24 Speaker Recognition Evaluation
NIST SRE 2024
Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification
ICNLSP 2024. ACL Anthology
Empirical approaches to variation: the case of the Timok variety of Torlak
Doctoral dissertation, University of Zurich
The LiRI Corpus Platform
CLARIN2023: Selected papers. Leuven: CLARIN
A Corpus-Based Analysis of the Grammatical Status of Short Demonstratives in the Timok Dialect
Journal of Slavic Linguistics 31(1–2), 245–269
Toward Sociolinguistic Corpora of Torlak
Zeitschrift für Slavische Philologie 79(1), 123–151
Under the Magnifying Glass: Dimensions of Variation in the Contemporary Timok Variety
Zeitschrift für Slavische Philologie 79(1), 153–194
Degrees of non-standardness: Feature-based analysis of variation in a Torlak dialect corpus
International Journal of Corpus Linguistics 27(2), 220–247
MultextEast V6 Torlak Specifications
In Tomaž Erjavec (ed.), MultextEast V6. CLARIN.SI
Representing variation in a spoken corpus of an endangered dialect: the case of Torlak
Language Resources and Evaluation. Springer Nature
Corpus-based analysis of spoken narratives: introducing a corpus and a search tool
Govor 37(2), 149–178. Zagreb: Hrvatsko filološko društvo
Acta Linguistica Petropolitana 16(2), 160–180 (in Russian)
Corpora and Processing Tools for Non-Standard Contemporary and Diachronic Balkan Slavic
Proceedings of the Student Research Workshop (RANLPStud 2019), RANLP 2019, Varna, 62–68
Prostorna raspodela frekvencije post-pozitivnog člana u timočkom govoru [Areal distribution of the frequency of the post-posed article in the Timok vernacular]
In S. Ćirković (ed.), Timok: folkloristička i lingvistička terenska istraživanja 2015–2017, 181–199
Creation and some ideas for classroom use of an electronic corpus of the dialect of Bunjevci
In J. Filipović & J. Vučo (eds.), Minority Languages in Education and Language Learning. Belgrade: Faculty of Philology
Digitalna zaštita bunjevačkih govora [Digital protection of the Bunjevac dialect]
In D. Njegovan (ed.), Kultura i identitet Bunjevaca. Novi Sad: Muzej Vojvodine, 473–489
MA thesis. Belgrade: Faculty of Philology
Interested in collaboration or consulting? Let's connect.