<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="http://korbinianfriedl.net/feed.xml" rel="self" type="application/atom+xml" /><link href="http://korbinianfriedl.net/" rel="alternate" type="text/html" /><updated>2026-07-08T10:58:21+00:00</updated><id>http://korbinianfriedl.net/feed.xml</id><title type="html">Korbinian Friedl</title><subtitle>Korbinian Friedl: Philosophy and A.I. research</subtitle><entry><title type="html">The Impossibility of Eliciting Latent Knowledge</title><link href="http://korbinianfriedl.net/research/2026/07/01/the-impossibility-of-eliciting-latent-knowledge.html" rel="alternate" type="text/html" title="The Impossibility of Eliciting Latent Knowledge" /><published>2026-07-01T00:00:00+00:00</published><updated>2026-07-01T00:00:00+00:00</updated><id>http://korbinianfriedl.net/research/2026/07/01/the-impossibility-of-eliciting-latent-knowledge</id><content type="html" xml:base="http://korbinianfriedl.net/research/2026/07/01/the-impossibility-of-eliciting-latent-knowledge.html"><![CDATA[<p><strong>Korbinian Friedl</strong>, Francis Rhys Ward, Paul Yushin Rapoport, Tom Everitt, Jonathan Richens</p>

<h2 id="abstract">Abstract</h2>
<p>Advanced AI systems have extensive knowledge of their environments; in fact, their knowledge may (far) exceed that of their developers or users. Consequently, a desirable property for an AI system is that it is honest – that it accurately reports its beliefs about the world. Designing an AI system to be honest may be difficult, especially if we want to ask it questions about latent variables in the environment – variables which are hidden from the human interacting with it. This gives rise to the problem of eliciting latent knowledge (ELK): the problem of training an AI agent to honestly report its beliefs. In this paper, we make ELK formally precise using Causal Influence Diagrams (CIDs). CIDs can be used to describe the relationship between an agent’s training environment and its subjective representation of the world. We use CIDs to formalise the distinction between observable and latent variables, to specify what exactly it means for an agent to be honest, and to formally define goal misgeneralisation. We show that, under certain circumstances, developers can incentivise an agent to honestly answer questions by providing correct feedback during training. However, a natural, but undesirable, way for an agent to generalise is to provide answers which humans would evaluate as true, rather than honest answers. We prove an impossibility theorem stating: There is no feedback-based training strategy that depends only on agent behaviour and with certainty produces an honest agent, even if feedback is perfect during training.</p>

<p><a href="https://arxiv.org/abs/2606.12268v1">Paper (PDF)</a></p>]]></content><author><name></name></author><category term="Research" /><category term="AI-safety" /><category term="formal-epistemology" /><category term="causal-diagrams" /><summary type="html"><![CDATA[Korbinian Friedl, Francis Rhys Ward, Paul Yushin Rapoport, Tom Everitt, Jonathan Richens]]></summary></entry><entry><title type="html">“Uncertainty Aware Review Hallucination for Science Article Classification” in Findings of the Association for Computational Linguistics</title><link href="http://korbinianfriedl.net/research/2021/08/01/uncertainty-aware-review-hallucination.html" rel="alternate" type="text/html" title="“Uncertainty Aware Review Hallucination for Science Article Classification” in Findings of the Association for Computational Linguistics" /><published>2021-08-01T00:00:00+00:00</published><updated>2021-08-01T00:00:00+00:00</updated><id>http://korbinianfriedl.net/research/2021/08/01/uncertainty-aware-review-hallucination</id><content type="html" xml:base="http://korbinianfriedl.net/research/2021/08/01/uncertainty-aware-review-hallucination.html"><![CDATA[<p><strong>Korbinian Friedl</strong>, Georgios Rizos, Lukas Stappen, Madina Hasan, Lucia Specia, Thomas Hain, Björn Schuller</p>

<h2 id="abstract">Abstract</h2>
<p>The high subjectivity and costs inherent in peer reviewing have recently motivated the preliminary design of machine learning-based acceptance decision methods. However, such approaches are limited in that they:</p>
<ul>
  <li>do not explore the usage of both the reviewer and area chair recommendations,</li>
  <li>do not explicitly model subjectivity on a per submission basis, and</li>
  <li>are not applicable in realistic settings, by assuming that review texts are available at test time, when these are exactly the inputs that should be considered to be missing in this application.</li>
</ul>

<p>We propose to utilise methods that
model the aleatory uncertainty of the submissions, while also exploring different loss importance interpolations between area chair and reviewers’ recommendations. We also propose a modality hallucination approach to impute review representations at test time, providing the first realistic evaluation framework for this challenging task.</p>

<p><em>Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021</em></p>

<p><a href="https://aclanthology.org/2021.findings-acl.443.pdf">Paper (PDF)</a> · <a href="https://aclanthology.org/2021.findings-acl.443/">ACL Anthology</a> · <a href="https://aclanthology.org/2021.findings-acl.443.bib">BibTeX</a></p>]]></content><author><name></name></author><category term="Research" /><category term="NLP" /><category term="machine-learning" /><category term="uncertainty" /><summary type="html"><![CDATA[Korbinian Friedl, Georgios Rizos, Lukas Stappen, Madina Hasan, Lucia Specia, Thomas Hain, Björn Schuller]]></summary></entry></feed>