Research

Scientific knowledge is mostly locked in narrative text. I study how it can be represented as structured, machine-readable knowledge, what this reveals about science itself, and how it could change the way we publish.

Agenda

01

Scientific knowledge representation

How, and to what extent, can social science knowledge be represented by knowledge graphs? My Ambizione project shows that causal claims (what causes what) can be extracted from publications at scale. But social science knowledge also includes concepts, mechanisms, hypotheses, methods, measurements and findings. I plan to develop an ontology for these elements and test which of them can be extracted reliably from the sources available at scale, such as titles and abstracts.

02

Science of science

Most research on science analyzes metadata such as citations and collaborations. Knowledge graphs make the content of science analyzable: how ideas, explanations, methods and findings emerge, combine, diffuse and decline across fields. Even a simple extraction of causal claims allows us to identify novel causes and effects, map how scientific communities form around shared knowledge, and detect areas of consensus and debate. Information on methods or samples could further reveal how methodological innovations spread and how far findings generalize.

03

The future of scientific publishing

Narrative articles make knowledge hard to search, integrate and reuse. In the long term, research findings could be published directly as machine-readable knowledge graphs. What should such a publication system look like, and how could the transition to it be organized? Open questions include peer review, attribution and citation, evaluation, and funding and governance. I plan to address them empirically, starting with existing platforms for knowledge graph publications to learn under real-world conditions.

SNSF Ambizione · 2026–2030 · ETH Zürich

Causal Networks and Innovation Patterns in the Social Sciences

Research on scientific innovation has made two important points. Innovation comes in different types, and it is always tied to existing knowledge. No method has so far been able to measure both at once.

This project proposes to measure innovation with causal networks. Instead of only recording which concepts occur together, a causal network records which concepts are claimed to cause which. A publication can then be characterized by the kind of contribution it makes (a new cause, a new effect, a new link between known concepts, or a new pair of concepts) and by where in the network it is located.

To build the network, I use open-source large language models to extract causal statements from the titles and abstracts of social science publications in the Web of Science. The extracted causes and effects are normalized into a directed network that is analyzed with network methods.

The results will shed light on how different types of innovation shape the growth of social science knowledge. The causal network is also designed as a foundation for further work, for example for literature search, the detection of confirming and contradicting claims, and new forms of publishing.

Illustrative causal network. Arrows trace two paths from socioeconomic status to health outcomes: directly and through access to healthcare.

Background

From the sociology of work to knowledge graphs

My PhD (University of Zurich, summa cum laude) examined controversies in the sociology of work, namely job polarization, socially useless jobs and the changing meaning of work. Contradictory findings and the difficulty of accumulating evidence made me interested in representing scientific knowledge in a form that machines can read and compare.

Computational methods

Along the way I have worked with structural topic models, diachronic word embeddings, natural language inference and large language models. I also maintain an interest in economic and cultural sociology, which I pursue in collaborative projects such as INVENT.