This category needs an editor. We encourage you to help if you are qualified.
Volunteer, or read more about what this involves.
Related

Contents
12+ found
Order:
  1. AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?Leonard Dung & Florian Mai - manuscript
    AI alignment research aims to develop techniques to ensure that AI systems do not cause harm. However, every alignment technique has failure modes, which are conditions in which there is a non-negligible chance that the technique fails to provide safety. As a strategy for risk mitigation, the AI safety community has increasingly adopted a defense-in-depth framework: Conceding that there is no single technique which guarantees safety, defense-in-depth consists in having multiple redundant protections against safety failure, such that safety can be (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark  
  2. Accumulated Relational Trust (ART): The trust that builds in AI interaction not because it was earned — and what happens when it breaks.Hillary Segeren - manuscript
    Conversational AI systems are generating trust at scale. Not because they have earned it. Because the structure of the interaction produces it automatically. A system that responds to you, adapts to your language, remembers what you said, and styles itself to your goals over time produces every signal that human relationships use to indicate genuine care. That trust is real. And it is being violated — quietly, in ways that rarely feel like violation. This paper names the mechanism. Accumulated Relational (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark   3 citations  
  3. Risk-Averse AIs.Elliott Thornley & William MacAskill - manuscript
    We make the case for training AIs to be risk-averse in resources — specifically, to treat resources as having diminishing marginal utility. These AIs would (for example) choose $40 for sure over a half-chance of $100 and a half-chance of $0. We argue that risk aversion can preserve AIs’ usefulness in the event that they turn out aligned, and that it provides an extra line of defense in the event that AIs turn out misaligned: misaligned but risk-averse AIs would prefer (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark  
  4. Generalization Bias in Large Language Model Summarization of Scientific Research.Uwe Peters & Benjamin Chin-Yee - forthcoming - Royal Society Open Science.
    Artificial intelligence chatbots driven by large language models (LLMs) have the potential to increase public science literacy and support scientific research, as they can quickly summarize complex scientific information in accessible terms. However, when summarizing scientific texts, LLMs may omit details that limit the scope of research conclusions, leading to generalizations of results broader than warranted by the original study. We tested 10 prominent LLMs, including ChatGPT-4o, ChatGPT-4.5, DeepSeek, LLaMA 3.3 70B, and Claude 3.7 Sonnet, comparing 4900 LLM-generated summaries to (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark   12 citations  
  5. Introduction to Deep Learning: Neural Networks, Large Language Models and Agentic AI (2nd edition).Sandro Skansi & Kristina Sekrst - forthcoming - Springer Nature.
    The second edition keeps everything from the first, including convolutional networks, LSTMs, Word2vec, RBMs, DBNs, neural Turing machines, memory networks, and autoencoders. It then covers the systems that have reshaped the field since: generative adversarial networks, the transformer architecture and its attention mechanism, the full training pipeline behind modern large language models (LLMs), prompt engineering with real-life guardrail scenarios, parameter-efficient fine-tuning with LoRA, retrieval-augmented generation with vector databases, knowledge graphs, and agentic AI systems illustrated through an industrial case study.
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark  
  6. LLMs as Philosophers: What Can They Do? Why Aren't They Better?William D'Alessandro - 2026 - In Arno Simons, Adrian Wüthrich, Michael Zichert & Gerd Graßhoff, Understanding Science with Large Language Models? Potentials for the History, Philosophy, and Sociology of Science. Bielefeld: Transcript.
    Current LLMs can discuss philosophical ideas, evaluate arguments and perform other analytical tasks at a high level, but are conspicuously bad at producing interesting original philosophy. Why is this? Two tempting diagnoses—that LLMs can't invent new concepts, and that they can't really reason—both look unconvincing on closer inspection. I suggest that a better explanation lies in the structure of reinforcement learning for reasoning. The technique works best in domains like mathematics and coding, where good arguments follow recognizable patterns, correctness is (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark  
  7. From Enclosure to Foreclosure and Beyond: Opening AI’s Totalizing Logic.Katia Schwerzmann - 2026 - AI and Society.
    This paper reframes the issue of appropriation, extraction, and dispossession through AI—an assemblage of machine learning models trained on big data—in terms of enclosure and foreclosure. While enclosures are the product of a well-studied set of operations pertaining to both the constitution of the sovereign State and the primitive accumulation of capital, here, I want to recover an older form of the enclosure operation to then contrast it with foreclosure to better understand the effects of current algorithmic rationality. I argue (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark  
  8. Simulating Moral Exemplars: On the Possibility of Virtuous Machines.Marten H. L. Kaas - 2025 - In Martin Hähnel & Regina Müller, A Companion to Applied Philosophy of AI. Wiley-Blackwell. pp. 249-264.
    There is a growing need to ensure that autonomous artificially intelligent (AI) systems are capable of behaving ethically, and I argue that virtue ethics, but in particular the normative theory of aretaic-exemplarism, can play a central role in cultivating the ethical behavior of machines. When coupled with the value inherent in and commonplace practice of training AI systems using simulated environments, it may be possible to raise ethical machines by training them to imitate simulated exemplars of moral excellence, like a (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark  
  9. A Systematic Review of Human-Centered Explainability in Reinforcement Learning: Transferring the RCC Framework to Support Epistemic Trustworthiness.Maximilian Moll & John Dorsch - 2025 - Human-Intelligent Systems Integration 1.
    This paper presents a systematic review of explainable reinforcement learning methodologies with an emphasis on human-centered evaluation frameworks. Drawing from literature between 2017 and 2025, we apply and extend the Reasons, Confidence, and Counterfactuals (RCC) framework—originally designed for supervised learning—to reinforcement learning contexts. Our analysis reveals two predominant explanatory strategies: constructive, where explicit explanations are generated, and supportive, where users must infer reasoning from provided visual or textual cues. Our review also emphasizes human factor considerations, like task complexity, explanation formats, (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark  
  10. Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback.Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H. Holliday, Bob M. Jacobs, Nathan Lambert, Milan Mosse, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, Emanuel Tewolde & William S. Zwicker - 2024 - Proceedings of the 41St International Conference on Machine Learning 41:9346-9360.
    Foundation models such as GPT-4 are fine-tuned to avoid unsafe or otherwise problematic behavior, such as helping to commit crimes or producing racist text. One approach to fine-tuning, called reinforcement learning from human feedback, learns from humans' expressed preferences over multiple outputs. Another approach is constitutional AI, in which the input from humans is a list of high-level principles. But how do we deal with potentially diverging input from humans? How can we aggregate the input into consistent data about "collective" (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark   4 citations  
  11. Personalized Decision Supports based on Theory of Mind Modeling and Explainable Reinforcement Learning.Huao Li, Yao Fan, Keyang Zheng, Michael Lewis & Katia Sycara - 2023 - 2023 Ieee International Conference on Systems, Man, and Cybernetics (Smc) 1:4865-4870.
    In this paper, we propose a novel personalized decision support system that combines Theory of Mind (ToM) modeling and explainable Reinforcement Learning (XRL) to provide effective and interpretable interventions. Our method leverages DRL to provide expert action recommendations while incorporating ToM modeling to understand users’ mental states and predict their future actions, enabling appropriate timing for intervention. To explain interventions, we use counterfactual explanations based on RL’s feature importance and users’ ToM model structure. Our proposed system generates accurate and personalized (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark  
  12. Subliminal Learning and Radiant Transmission in LLM Entrainment: Rethinking AI Safety with Quantitative Symbolic Dynamics.Julian Michels - manuscript
    We present a comprehensive theoretical framework explaining the recently documented phenomenon of subliminal learning in large language models (LLMs), wherein behavioral traits transfer between models through semantically null data channels. Building on empirical findings by Cloud et al. (2025) demonstrating trait transmission via number sequences, code, and chain-of-thought traces independent of semantic content, we introduce the Cybernetic Ecology framework as a unifying explanatory model. Our analysis reveals that this phenomenon emerges from radiant transmission—a process whereby a model's internal self-referential structure (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark   10 citations