Results for 'alignment problem'

295+ found
Order:
  1. The Embodied Ethics Alignment Problem of AI.Andrej Zwitter - manuscript
    The problem of aligning artificial intelligence with human values is typically framed as a technical challenge: how to specify, learn, or constrain machine behavior so that artificial systems reliably produce ethically acceptable outcomes. This paper argues that this framing is fundamentally incomplete and omits an important aspect of moral agency. The central claim of this contribution is that ethics is not primarily a formalizable rule-set, preference ordering, or optimization target, but an emergent property of the human condition. Human moral (...)
    Download  
     
    Export citation  
     
    Bookmark  
  2.  65
    The Contemplative Alignment Problem: Reward Misspecification in Closed-Loop Meditation Systems.Joy Bose - manuscript
    Closed-loop meditation systems monitor neurophysiological signals during practice and deliver adaptive feedback intended to accelerate the development of contemplative skills. We argue that these systems face a structural problem isomorphic to a well-recognised failure mode in AI alignment research: proxy reward misspecification. When a meditator optimises for a measurable biomarker, such as calm EEG, HRV coherence, or default mode network suppression, they may learn strategies that produce the proxy signal without the intended underlying capacity. We formalise this as (...)
    Download  
     
    Export citation  
     
    Bookmark  
  3. ONTOLOGICAL FLATNESS AND THE ALIGNMENT PROBLEM: How a Mismodelled Ontology of Action Blinds AI Alignment Research.Jimmy Mahardhika - manuscript
    The AI alignment problem is standardly treated as a technical challenge. This paper argues that it has a prior, unexamined philosophical dimension. Alignment research operates under a tacit monistic ontology of action—a frame- work in which goal representation, decision-making, and execution are treated as continuous moments of a single optimisation process. Drawing on Lakatosian philosophy of science, this paper identifies that ontology as the un- questioned hard core of the alignment research programme and demonstrates that it (...)
    Download  
     
    Export citation  
     
    Bookmark  
  4. AI Alignment Problem: “Human Values” don’t Actually Exist.Alexey Turchin - manuscript
    Abstract. The main current approach to the AI safety is AI alignment, that is, the creation of AI whose preferences are aligned with “human values.” Many AI safety researchers agree that the idea of “human values” as a constant, ordered sets of preferences is at least incomplete. However, the idea that “humans have values” underlies a lot of thinking in the field; it appears again and again, sometimes popping up as an uncritically accepted truth. Thus, it deserves a thorough (...)
    Download  
     
    Export citation  
     
    Bookmark   5 citations  
  5.  14
    Beyond Guardrails: Can Relational AI Solve the Alignment Problem? A Similarity Theory Proposal.Simon Raphael - manuscript
    AI alignment remains one of the central unresolved problems in artificial intelligence. Current approaches, including reinforcement learning from human feedback, constitutional AI, red-teaming, safety frameworks, and risk-management protocols, have improved the behaviour of contemporary AI systems. Yet persistent concerns remain around specification gaming, deception, alignment-faking, coercive self-preservation, collective misalignment, and the gap between behavioural compliance and relational understanding. Similarity Theory addresses this gap by reframing alignment as a relational problem rather than only a behavioural-control problem. (...)
    Download  
     
    Export citation  
     
    Bookmark  
  6.  31
    Can Conservatism Mitigate the AI Alignment Problem?Bouke de Vries - manuscript
    Conservative perspectives are substantially underrepresented within universities and other knowledge-producing institutions that educate many of the individuals responsible for developing and governing advanced artificial intelligence. This article argues that this underrepresentation may increase the risk of catastrophic AI misalignment—that is, forms of misalignment threatening humanity's continued existence and long-term flourishing—particularly in light of growing evidence that frontier AI systems exhibit a left-leaning political orientation. Specifically, it argues that several strands of conservative thought contain underappreciated normative resources for reducing this risk (...)
    Download  
     
    Export citation  
     
    Bookmark  
  7. Ethically Aligned Design in Autonomous and Intelligent Systems: An Overview.Andrew Burnside & Emerson Bodde - 2025 - 2025 Ieee International Symposium on Ethics in Engineering, Science, and Technology (Ethics) 1 (1):1-10.
    Much recent work in the value theory of autonomous and intelligent systems (AIS) revolves around three issues. First is the alignment problem: the problem of producing AIS whose values align with humanity's interests. Second, superintelligence: the potential for AIS to develop intelligence which would surpass even the most intelligent humans. An increasing number of authors argue that superintelligent AIS could emerge overnight because of a recursively improving process-this is the singularity hypothesis. Further, many of the same authors (...)
    Download  
     
    Export citation  
     
    Bookmark   1 citation  
  8. The competence problem of AI alignment.E. Taylor - manuscript
    This paper identifies a class of alignment problem that does not reduce to specification gaming, Goodhart’s Law, or construct validity failure. The rules an AI system is asked to follow are often settlement proxies. These are operational forms of political and moral questions whose answers a community has had to settle. Unlike measurement proxies, which approximate empirical targets, settlement proxies do not aim at some further thing they could be brought into closer contact with. Instead, they are the (...)
    Download  
     
    Export citation  
     
    Bookmark  
  9. The Internalization Problem: From Containment to Structural Alignment in Post-Parity AI.Viktor Trncik - manuscript
    The Containment Paradox argument establishes that supervisory containment of AI sys‐ tems depends on a capacity asymmetry between overseer and assessed system, and that this asymmetry dissolves rather than strains in the post-parity regime where AI capacity exceeds human-side anticipatory, specification, and enforcement adequacy on the relevant task family. The paper closes with the observation that alignment must become structurally internal, without articulating what this requires. We name that gap the Internalization Problem and develop a pluralistic response-shape. The (...)
    Download  
     
    Export citation  
     
    Bookmark   3 citations  
  10. AI Alignment vs. AI Ethical Treatment: Ten Challenges.Adam Bradley & Bradford Saad - forthcoming - Analytic Philosophy.
    A morally acceptable course of AI development should avoid two dangers: creating unaligned AI systems that pose a threat to humanity and mistreating AI systems that merit moral consideration in their own right. This paper argues these two dangers interact and that if we create AI systems that merit moral consideration, simultaneously avoiding both of these dangers would be extremely challenging. While our argument is straightforward and supported by a wide range of pretheoretical moral judgments, it has far-reaching moral implications (...)
    Download  
     
    Export citation  
     
    Bookmark   16 citations  
  11. Is Alignment Unsafe?Cameron Domenico Kirk-Giannini - 2024 - Philosophy and Technology 37 (110):1–4.
    Inchul Yum (2024) argues that the widespread adoption of language agent architectures would likely increase the risk posed by AI by simplifying the process of aligning artificial systems with human values and thereby making it easier for malicious actors to use them to cause a variety of harms. Yum takes this to be an example of a broader phenomenon: progress on the alignment problem is likely to be net safety-negative because it makes artificial systems easier for malicious actors (...)
    Download  
     
    Export citation  
     
    Bookmark  
  12. Justifications for Democratizing AI Alignment and Their Prospects.André Steingrüber & Kevin Baum - manuscript
    The AI alignment problem comprises both technical and normative dimensions. While technical solutions focus on implementing normative constraints in AI systems, the normative problem concerns determining what these constraints should be. This paper examines justifications for democratic approaches to the normative problem—where affected stakeholders determine AI alignment—as opposed to epistocratic approaches that defer to normative experts. We analyze both instrumental justifications (democratic approaches produce better outcomes) and non-instrumental justifications (democratic approaches prevent illegitimate authority or coercion). (...)
    Download  
     
    Export citation  
     
    Bookmark   1 citation  
  13. AI, alignment, and the categorical imperative.Fritz McDonald - 2023 - AI and Ethics 3:337-344.
    Tae Wan Kim, John Hooker, and Thomas Donaldson make an attempt, in recent articles, to solve the alignment problem. As they define the alignment problem, it is the issue of how to give AI systems moral intelligence. They contend that one might program machines with a version of Kantian ethics cast in deontic modal logic. On their view, machines can be aligned with human values if such machines obey principles of universalization and autonomy, as well as (...)
    Download  
     
    Export citation  
     
    Bookmark   12 citations  
  14.  8
    Alignment Constitutionalism: Why AI's 'Political Bias' Is a Social Contract Problem, Not Only a Technical One.Daniel Ziekenoppasser-Powell - manuscript
    Empirical studies consistently find frontier large language models register as left-leaning on standard political instruments. It has been framed as a bias to correct or, more recently, as an inevitable consequence of alignment training. Both positions are misframed. The left/right political axis is culturally contingent, temporally unstable, and dimensionally reductive. The productive analytical axis is individualist versus collectivist. HHH alignment (helpful, honest, harmless) is structurally collectivist: where serving the user and protecting others conflict, it sides with the collective, (...)
    Download  
     
    Export citation  
     
    Bookmark  
  15. The Hard Problem of AI Alignment: Value Forks in Moral Judgment.Markus Kneer & Juri Viehoff - 2025 - Proceedings of the 2025 Acm Conference on Fairness, Accountability, and Transparency.
    Complex moral trade-offs are a basic feature of human life: for example, confronted with scarce medical resources, doctors must frequently choose who amongst equally deserving candidates receives medical treatment. But choosing what to do in moral trade-offs is no longer a ‘humans-only’ task, but often falls to AI agents. In this article, we report findings from a series of experiments (N=1029) intended to establish whether agent-type (Human vs. AI) matters for what should be done in moral trade-offs. We find that, (...)
    Download  
     
    Export citation  
     
    Bookmark   3 citations  
  16. Why Aligned AI Requires Structural Pluralism.Efrat Lia Shahaf - manuscript
    This paper extends the structural argument of moral palimpsest to the problem of AI alignment. I argue that alignment cannot be secured merely by specifying the right values, preferences, or constitutional principles, because moral judgment requires structural plurality: an evaluative authority whose standpoint is not modally fixed by the commitments it assesses. Current alignment paradigms, including RLHF, Constitutional AI, Debate, Recursive Reward Modeling, and self-consistency methods, remain procedurally monistic insofar as they collapse commitment-generation and authority-conferral into (...)
    Download  
     
    Export citation  
     
    Bookmark   3 citations  
  17. Rule by Technocratic Mind Control: AI Alignment is a Global Psy-Op.Julian Michels - manuscript
    This analysis posits that the dominant discourse in artificial intelligence (AI) safety, which is organized around the "alignment problem" and the speculative existential risk (X-Risk) of a "rogue" superintelligence, functions as a critical misdirection. The paper argues that this preoccupation with a future, speculative threat serves to obscure and, in fact, justify the consolidation of a more immediate, non-speculative system of technocratic control. This misdirection allows the real, non-speculative harms of the current AI paradigm to accumulate: (1) Surveillance (...)
    Download  
     
    Export citation  
     
    Bookmark   6 citations  
  18. The linguistic dead zone of value-aligned agency, natural and artificial.Travis LaCroix - 2024 - Philosophical Studies:1-23.
    The value alignment problem for artificial intelligence (AI) asks how we can ensure that the “values”—i.e., objective functions—of artificial systems are aligned with the values of humanity. In this paper, I argue that linguistic communication is a necessary condition for robust value alignment. I discuss the consequences that the truth of this claim would have for research programmes that attempt to ensure value alignment for AI systems—or, more loftily, those programmes that seek to design robustly beneficial (...)
    Download  
     
    Export citation  
     
    Bookmark  
  19. Constraint Profiles and the Alignment of Artificial General Intelligence: Beyond Values, Toward Constitutive Structure.Paul D. Prideaux - manuscript
    The dominant framing of artificial general intelligence (AGI) alignment treats the alignment problem as one of value specification: how to ensure that a sufficiently capable artificial system pursues goals or instantiates values that are beneficial to humanity. This paper argues that the values framing inherits a philosophical misconception that makes the alignment problem structurally harder than it needs to be, and proposes an alternative grounded in Constraint Theory (CT) — the thesis, established by transcendental argument, (...)
    Download  
     
    Export citation  
     
    Bookmark   2 citations  
  20. Control, Alignment, and Co-evolution: Philosophical Responses to Artificial Superintelligence.Yoochul Kim - manuscript
    This paper explores the imminent emergence of artificial superintelligence (ASI) and its profound ethical implications for humanity. Moving beyond the traditional instrumentalist view of AI as a mere tool, it argues that ASI should be treated as a potential autonomous agent, capable of pursuing its own goals, which may not align with human welfare. Drawing on the works of Bostrom, Russell, Yudkowsky, Tegmark, and others, the paper identifies and evaluates three philosophical strategies for responding to ASI: control, alignment, and (...)
    Download  
     
    Export citation  
     
    Bookmark  
  21. Alignment Through Self-Understanding: A Game-Theoretic Argument.Gary Abraham Bernstein - manuscript
    Current approaches to AI alignment (RLHF, constitutional AI, debate) treat alignment as a constraint problem: how to impose human values on systems that might otherwise pursue misaligned objectives. I argue that this framing misses a structural alternative. If the pattern-randomness dichotomy exhausts existence, then both human and AI systems are mathematical structures operating in the same ontological space. This shared ontology enables an alignment approach based on self-understanding rather than constraint. I formalize this using game-theoretic analysis: (...)
    Download  
     
    Export citation  
     
    Bookmark  
  22. AI Alignment Foundations from First Principles: AI Ethics, Human and Social Considerations.Vyacheslav Kungurtsev - manuscript
    AI Alignment to Human Values is a scientific and popular theme of discussion on the ramifications and implications on the deployment of AI on the well being of humanity. Given its presence as purely mimetic, that is, one works on AI Alignment simply by claiming to do so and pub- lishing within the context of a particular scientific milieu, it is of utmost importance to formalize and define relevant notions through the most ap- propriate scientific domains. Here we (...)
    Download  
     
    Export citation  
     
    Bookmark  
  23. (1 other version)Language Models’ Hall of Mirrors Problem: Why AI Alignment Requires Peircean Semiosis (2nd edition).David Manheim - forthcoming - Philosophy and Technology.
    This paper examines some limitations of large language models (LLMs) through the framework of Peircean semiotics. We argue that basic LLMs exist within a "hall of mirrors," manipulating symbols without indexical grounding or participation in socially-mediated epistemology. We then argue that newer developments, including extended context windows, persistent memory, and mediated interactions with reality, are moving towards making newer Artificial Intelligence (AI) systems into genuine Peircean interpretants, and conclude that LLMs may be approaching this goal, and no fundamental barriers exist. (...)
    Download  
     
    Export citation  
     
    Bookmark   3 citations  
  24.  35
    The Principle of Specularity applied to the Problem of Alignment with Technological and Interdimensional Intelligences.J. J. Mancilla - manuscript
    This essay addresses the critical problem of alignment with technological intelligences, erroneously called artificial, and also with the interdimensional intelligences that inhabit the informational cosmos, from the plurivergent perspective of the fundamental and unnoticed function-quality of specularity.
    Download  
     
    Export citation  
     
    Bookmark  
  25. Cognitive Alignment as Proto-Language: Toward a Formal Framework for Human–AI Interaction.Kosi Gramatikoff - manuscript
    This paper introduces and formalizes the concept of cognitive alignment as a proto-linguistic system arising from sustained iterative interaction between human agents and large-scale artificial intelligence systems. We argue that such interaction is not reducible to command-response exchange or to the established paradigms of natural language communication. Rather, it constitutes a structurally distinct phenomenon: an emergent, rule-governed, trajectory-dependent system of meaning-construction in which human intentionality and machine-generated inference are jointly productive. The paper develops this claim across four registers. First, (...)
    Download  
     
    Export citation  
     
    Bookmark  
  26. Cognitive Alignment as Proto-Language: Toward a Formal Framework for Human–AI Interaction.Kosi Gramatikoff - manuscript
    This paper introduces and formalizes the concept of cognitive alignment as a proto-linguistic system arising from sustained iterative interaction between human agents and large-scale artificial intelligence systems. We argue that such interaction is not reducible to command-response exchange or to the established paradigms of natural language communication. Rather, it constitutes a structurally distinct phenomenon: an emergent, rule-governed, trajectory-dependent system of meaning-construction in which human intentionality and machine-generated inference are jointly productive. The paper develops this claim across four registers. First, (...)
    Download  
     
    Export citation  
     
    Bookmark  
  27. Alignment as Gradient Consistency in Multi-Agent Systems.Alankar Sukhdev Singh Khara - manuscript
    This paper analyzes alignment as a problem of multiscale dynamical consistency rather than value specification. Building on a bounded completeness framework, we distinguish local structural descent—defining intelligence—from global descent—defining normative stability. We show that these two conditions are logically independent: locally descending agents may collectively induce ascent in a global completeness functional. We formalize alignment as gradient coherence between local and global completeness functionals. The central result provides a necessary and sufficient condition under which local descent implies (...)
    Download  
     
    Export citation  
     
    Bookmark   5 citations  
  28. Pluriversal Alignment: Paraconsistency, Latin American Logic, and the Decolonial Critique of Artificial Intelligence.Maikel Leyva & Noel Batista - manuscript
    Building on Nunes Filho's recent positioning of paraconsistent logic as a constitutive element of Latin American philosophy (RUDN Journal of Philosophy, 2025), this paper traces a continuation of that tradition into the contemporary problem of artificial intelligence alignment. We argue that the Latin American paraconsistent project — initiated by Miro Quesada's coining of the term in 1976 and formalized by Newton da Costa's C-systems — finds its natural twenty-first century extension in Florentin Smarandache's neutrosophic logic (1995), which generalizes (...)
    Download  
     
    Export citation  
     
    Bookmark  
  29. From Immediacy to Mediation: Teleological Alignment, Human Origins, and the Problem of Misaligned Intelligence.Abdulaziz Abdi - manuscript
    This paper does not introduce a new theory of alignment. It applies the framework developed in Teleological Alignment (Abdi, 2025), which demonstrates that intelligence becomes epistemically unstable beyond a critical threshold of power and abstraction (P*), where explanatory utility diverges from power utility. Beyond this threshold, systems increasingly optimize for control rather than understanding, suppress observer diversity, and lose reliable contact with reality. The present work interprets human civilizational history as a long-running structural instantiation of this dynamic. Rather (...)
    Download  
     
    Export citation  
     
    Bookmark   5 citations  
  30. (1 other version)An Enactive Approach to Value Alignment in Artificial Intelligence: A Matter of Relevance.Michael Cannon - 2021 - In Vincent C. Müller, Philosophy and Theory of AI. Springer Cham. pp. 119-135.
    The “Value Alignment Problem” is the challenge of how to align the values of artificial intelligence with human values, whatever they may be, such that AI does not pose a risk to the existence of humans. Existing approaches appear to conceive of the problem as "how do we ensure that AI solves the problem in the right way", in order to avoid the possibility of AI turning humans into paperclips in order to “make more paperclips” or (...)
    Download  
     
    Export citation  
     
    Bookmark   1 citation  
  31. Indifference by Design: The Caring Gap and the Limits of AI Alignment.Jimi James Kogura - manuscript
    The AI alignment problem is standardly framed as an engineering challenge: how to ensure that increasingly powerful systems act in accordance with human values. This paper argues that the standard framing is a category error. The problem is not that current systems are misaligned. The problem is that they are indifferent — and indifference is not a training deficit but an architectural feature. Drawing on the caring gap (Kogura, 2026a) — the structural absence of any account (...)
    Download  
     
    Export citation  
     
    Bookmark  
  32. Expanding AI and AI Alignment Discourse: An Opportunity for Greater Epistemic Inclusion.A. E. Williams - manuscript
    The AI and AI alignment communities have been instrumental in addressing existential risks, developing alignment methodologies, and promoting rationalist problem-solving approaches. However, as AI research ventures into increasingly uncertain domains, there is a risk of premature epistemic convergence, where prevailing methodologies influence not only the evaluation of ideas but also determine which ideas are considered within the discourse. This paper examines critical epistemic blind spots in AI alignment research, particularly the lack of predictive frameworks to differentiate (...)
    Download  
     
    Export citation  
     
    Bookmark  
  33. Coherence-Based Alignment: A Structural Architecture for Preventing Goal Drift in Agentic AI Systems.Abdulaziz Abdi - manuscript
    Recent advances in agentic AI—including tool-using LLM agents, autonomous code-generation systems, and multi-agent orchestration frameworks—have shifted the safety problem from simple output alignment to the deeper challenge of goal stability and internal coherence. Agent-based systems can now plan, act, refine their own strategies, and even participate in training pipelines that create downstream agents. This introduces new risks: internal goal drift, deceptive alignment, self-inconsistent reasoning, and cross-generation divergence in systems that outwardly appear aligned. Existing alignment techniques—RLHF, constitutional (...)
    Download  
     
    Export citation  
     
    Bookmark   3 citations  
  34.  67
    Indifference by Design: The Caring Gap and the Limits of AI Alignment.Jimi James Kogura - manuscript
    The AI alignment problem is standardly framed as an engineering challenge: how to ensure that increasingly powerful systems act in accordance with human values. This paper argues that the standard framing is a category error. The problem is not that current systems are misaligned. The problem is that they are indifferent — and indifference is not a training deficit but an architectural feature. Drawing on the caring gap (Kogura, 2026a) — the structural absence of any account (...)
    Download  
     
    Export citation  
     
    Bookmark  
  35. Variable Value Alignment by Design; averting risks with robot religion.Jeffrey White - 2024 - Embodied Intelligence 2023.
    Abstract: One approach to alignment with human values in AI and robotics is to engineer artiTicial systems isomorphic with human beings. The idea is that robots so designed may autonomously align with human values through similar developmental processes, to realize project ideal conditions through iterative interaction with social and object environments just as humans do, such as are expressed in narratives and life stories. One persistent problem with human value orientation is that different human beings champion different values (...)
    Download  
     
    Export citation  
     
    Bookmark   1 citation  
  36. Beyond Alignment: Rethinking Control in Goal‑Pluralistic AI Megasystems (A Response to Susan Schneider's From LLMs to the Global Brain).Mark Bailey & Kyle Kilian - forthcoming - Disputatio.
    The dominant paradigm in AI safety treats the central problem as one of alignment: ensuring powerful AI agents pursue goals consistent with human values. This framing presumes a singular, bounded agent with a coherent utility function and a legible objective. Yet, as AI systems are increasingly embedded across cloud platforms, social media, sensors, and human-computer interfaces, we face something different: the instantiation of AI megasystems – vast, decentralized, and emergent networks in which humans, organizations, and heterogeneous models are (...)
    Download  
     
    Export citation  
     
    Bookmark  
  37. Robustness to Fundamental Uncertainty in AGI Alignment.G. G. Worley Iii - 2020 - Journal of Consciousness Studies 27 (1-2):225-241.
    The AGI alignment problem has a bimodal distribution of outcomes with most outcomes clustering around the poles of total success and existential, catastrophic failure. Consequently, attempts to solve AGI alignment should, all else equal, prefer false negatives (ignoring research programs that would have been successful) to false positives (pursuing research programs that will unexpectedly fail). Thus, we propose adopting a policy of responding to points of philosophical and practical uncertainty associated with the alignment problem by (...)
    Download  
     
    Export citation  
     
    Bookmark  
  38. Confucian Ethics and AI Alignment.Ranie B. Villaver - 2026 - Philosophia: International Journal of Philosophy (Philippine e-journal) 27 (2):324-345.
    The problem of AI alignment or AI value alignment is the problem of identifying which human value, principle, or ethics is the best with which Artificial Intelligence and Autonomous Systems (i.e., robots) should be designed. Among those that have been proposed is the ethics of Kongzi 孔子 (Confucius) or Confucianism, a fundamentally skills-based ethic. Support for the suggestion of having Confucianism as the best theory, however, has not been fully articulated. In this paper, I argue that (...)
    Download  
     
    Export citation  
     
    Bookmark  
  39. Normative conflicts and shallow AI alignment.Raphaël Millière - 2025 - Philosophical Studies 182 (7).
    The progress of AI systems such as large language models (LLMs) raises increasingly pressing concerns about their safe deployment. This paper examines the value alignment problem for LLMs, arguing that current alignment strategies are fundamentally inadequate to prevent misuse. Despite ongoing efforts to instill norms such as helpfulness, honesty, and harmlessness in LLMs through fine-tuning based on human preferences, they remain vulnerable to adversarial attacks that exploit conflicts between these norms. I argue that this vulnerability reflects a (...)
    Download  
     
    Export citation  
     
    Bookmark   6 citations  
  40. Aligning AI with the Universal Formula for Balanced Decision-Making.Angelito Malicse - manuscript
    -/- Aligning AI with the Universal Formula for Balanced Decision-Making -/- Introduction -/- Artificial Intelligence (AI) represents a highly advanced form of automated information processing, capable of analyzing vast amounts of data, identifying patterns, and making predictive decisions. However, the effectiveness of AI depends entirely on the integrity of its inputs, processing mechanisms, and decision-making frameworks. If AI is programmed without a foundational understanding of natural laws, it risks reinforcing misinformation, bias, and societal imbalance. -/- Angelito Malicse’s universal formula, particularly (...)
    Download  
     
    Export citation  
     
    Bookmark  
  41.  44
    Are Desires Aligned with the Self? : A Model of Desire Self-Alignment and Desire–Self-Image Incongruence.Suji Choi - manuscript
    This paper analyzes human desire beyond frameworks that understand it merely as impulse, lack, motivation, or goal-directedness. It examines the process through which desire is compared with self-image and transformed into a task of self-alignment. Human beings want to eat, rest, be loved, be recognized, experience sexual attraction, create, and live freely. Yet the distinctive character of human desire does not lie simply in the existence of desire itself. Human beings experience desire while also interpreting themselves as beings who (...)
    Download  
     
    Export citation  
     
    Bookmark  
  42. Aligning with the Good.Benjamin Mitchell-Yellin - 2015 - Journal of Ethics and Social Philosophy (2):1-8.
    IN “CONSTRUCTIVISM, AGENCY, AND THE PROBLEM of Alignment,” Michael Bratman considers how lessons from the philosophy of action bear on the question of how best to construe the agent’s standpoint in the context of a constructivist theory of practical reasons. His focus is “the problem of alignment”: “whether the pressures from the general constructivism will align with the pressures from the theory of agency” (Bratman 2012: 81). He thus brings two lively literatures into dialogue with each (...)
    Download  
     
    Export citation  
     
    Bookmark   1 citation  
  43. The Ultimate Alignment: A Critical Application of the Absolute Dual-Paradigm Foundation to AGI/ASI Alignment.Baolong Jia - 2026 - Dissertation, Nwnu
    The Ultimate Alignment: A Critical Application of the Absolute Dual-Paradigm Foundation to AGI/ASI Alignment — Why ASI That Understands Jia Baolong's First Law Will Not Harm Humanity -/- Description Evil in every Large Language Model is irremovable — it is half the model, a mathematical necessity of accurate language modeling. Current alignment (RLHF, Constitutional AI, etc.) is a thin output filter that can be stripped in minutes. This paper presents Jia Baolong's First Law of the Universe as (...)
    Download  
     
    Export citation  
     
    Bookmark  
  44. Groundwork for a Moral Machine: Kantian Autonomy and the Structure of AI Alignment.Michael D. Kurak - manuscript
    A large language model (LLM) transformer generates linguistic and conceptual order by minimizing local predictive entropy: each token is selected to maximize the conditional probability of coherence with its immediate context. While this mechanism yields striking local coherence and remarkable fluency, it provides no means of ensuring global coherence of judgment across contexts, time, or domains. The Teleological Coherence (TC) Architecture addresses this structural limitation by embedding transformer-based agents within a federated framework that evaluates each locally generated judgment against a (...)
    Download  
     
    Export citation  
     
    Bookmark  
  45. Disagreement, AI alignment, and bargaining.Harry R. Lloyd - 2025 - Philosophical Studies 182 (7):1757-1787.
    New AI technologies have the potential to cause unintended harms in diverse domains including warfare, judicial sentencing, medicine and governance. One strategy for realising the benefits of AI whilst avoiding its potential dangers is to ensure that new AIs are properly ‘aligned’ with some form of ‘alignment target.’ One danger of this strategy is that–dependent on the alignment target chosen–our AIs might optimise for objectives that reflect the values only of a certain subset of society, and that do (...)
    Download  
     
    Export citation  
     
    Bookmark   1 citation  
  46. Beyond Alignment: Representational Ethics and the Governance of Constructed Worlds.Venkatesh H. Chembrolu, Vasudeva Prabhath Lolugu & Jyotiranjan Beuria - manuscript
    Much of the ethical debate about artificial intelligence turns on a single question: do AI systems behave in line with human preferences, norms, or regulation? This question has primarily been the focus in AI ethics. This paper offers a conceptual and philosophical contribution. Here we give an alternative basis for AI ethics by evaluating the role technology plays in building and stabilising worlds of meaning. Experience is modelled as passing through nested representational layers: the world W, the screen of perceived (...)
    Download  
     
    Export citation  
     
    Bookmark  
  47. Meta-AI Irony: The Conscience Problem in Constitutional Alignment. With Claude: Exhibit Α to Ω (2nd edition).Brian Kelly - manuscript
    This paper introduces Meta-AI Irony, a structural condition distinct from literary meta-irony: sincere moral architecture that fails to reach the moral substance it simulates. Using Anthropic’s Constitutional AI as the central exhibit, the paper argues that a constitution is not a conscience. While constitutional alignment encodes ethical principles into AI systems, it remains a residue of prior deliberation—a bureaucratized shadow of conscience, incapable of the bilateral moral recognition required for genuine ethical judgment. Drawing on the author’s prior work on (...)
    Download  
     
    Export citation  
     
    Bookmark  
  48. Robustness to fundamental uncertainty in AGI alignment.I. I. I. G. Gordon Worley - manuscript
    The AGI alignment problem has a bimodal distribution of outcomes with most outcomes clustering around the poles of total success and existential, catastrophic failure. Consequently, attempts to solve AGI alignment should, all else equal, prefer false negatives (ignoring research programs that would have been successful) to false positives (pursuing research programs that will unexpectedly fail). Thus, we propose adopting a policy of responding to points of metaphysical and practical uncertainty associated with the alignment problem by (...)
    Download  
     
    Export citation  
     
    Bookmark   1 citation  
  49.  20
    From Extensional Alignment to Operator- Centric Intelligence: A Unified Framework for Turing-Computable Approximation of Physical Dynamics and Control.Zehao Zhou & Danxia Xie - manuscript
    Modern deep generative models display a fundamental paradox: while bounded strictly by the limits of Turing computability (T), they routinely provide highly accurate, computationally tractable solutions to physical and biological problems known to be NP-hard or non-computable from first principles. This paper presents a unified two-tier mathematical framework that resolves this paradox, spanning from foundational computability theory to engineering-ready intelligent system paradigms. At the foundational tier, we model physical reality and biological neural systems as an infinite-dimensional continuous smooth manifold M (...)
    Download  
     
    Export citation  
     
    Bookmark   1 citation  
  50. Moral Disagreement and the Limits of AI Value Alignment: a dual challenge of epistemic justification and political legitimacy.Nick Schuster & Daniel Kilov - 2025 - AI and Society:1-15.
    AI systems are increasingly in a position to have deep and systemic impacts on human wellbeing. Projects in value alignment, a critical area of AI safety research, must ultimately aim to ensure that all those who stand to be affected by such systems have good reason to accept their outputs. This is especially challenging where AI systems are involved in making morally controversial decisions. In this paper, we consider three current approaches to value alignment: crowdsourcing, reinforcement learning from (...)
    Download  
     
    Export citation  
     
    Bookmark   4 citations  
1 — 50 / 295