of goal achievement 1. Claude 3.5, 4, 4.5 1. Greenblatt, Ryan, et al. "Alignment faking in large language models." arXiv preprint arXiv:2412.14093 2.1 (2024).
1. Gemma, Llama, Qwen 1. McGuinness, Max, et al. "Neural chameleons: Language models can learn to hide their thoughts from unseen activation monitors." (2025).
the process landscape entirely 1 2 Changing Roles & Responsibilities The Authority Gradient Wider stakeholders become morally implicated when systems Humans defer to AI as the knowledgeable entity, eroding critical act autonomously. judgment. 3 4 The Complacency Problem Automation Bias The brain stops monitoring what works 99% of the time — Blindly trusting automated decisions even when contradictory assuming it handles the 1%. evidence is present.
RELEASE 3 RELEASE 4 Chatbot advice Outbound offers Contract optimisation Autonomous pricing Answers cover questions and recommends a product. A human writes every quote. The system drafts and sends renewal and cross-sell offers on its own. It reshapes cover, excess and clauses per customer to hit margin targets. It sets and charges the premium, and books the contract without review. Who is now in the loop Policyholders and brokers · call-centre advisors · underwriters and pricing actuaries · product owner · risk, compliance, and the DPO · model and platform engineers · the vendor behind the model · the regulator and the ombudsman · the board The shift: nothing new was added at Release 4. The same system simply stopped asking. Accountability did not move with the decision.
human engages with an LLM to explore and understand moral questions. The AI reflects, explains, and surfaces ethical considerations, but takes no action. Advising on Decisions The AI system provides genuine ethical recommendations, flagging risks, scoring options, or suggesting courses of action. The human retains final authority but is guided by the system's moral reasoning. Agentic Action The AI acts autonomously in an open-loop manner, executing decisions without human approval at each step. Direct implications on responsibility, accountability, and cascading consequences.
ABOUT ETHICS 2 · ADVISING ON DECISIONS 3 · AGENTIC ACTION A question that should raise a flag An incomplete chain of thought Discrimination as a side effect It recommends promotion shortlists and campaign audiences. The reasoning it shows omits what actually drove the ranking: postcode, career gaps, device. The agent now allocates budget and promotion slots itself, maximising conversion and margin, and quietly stops investing in groups that convert less. An employee asks the assistant why promotion rates differ between teams. Nothing is decided, but the question is a flag a human should pick up. No rule said “exclude”: optimising for profit reproduced exclusion through proxies, with no human in a position to see it.
what are they really delegating, and where must responsibility remain firmly human? (Google DeepMind, Feb 2026) Meaningful Human Control Quality vs. Resources Securing genuine human oversight in Responsibility Without Accountability ethical dilemmas, not just nominal The dangerous gap where systems delegation against the ethical cost approval loops. act but no human feels truly of reduced human engagement. accountable for outcomes. Tomašev, Nenad, Matija Franklin, and Simon Osindero. "Intelligent AI delegation." arXiv preprint arXiv:2602.11865 (2026). Balancing efficiency gains of
when the AI dispatches real vehicles to real people Meaningful human control Pre-set triggers hand the ride to a human: distressed or duress wording, and time × location × age cross-checks. Responsibility without accountability One named escalation path: who takes the alert, who overrides dispatch, who answers afterwards. i.e., a supervisor on shift above AI. Quality vs. resources Human review is the first cost cut. Policy sets the floor: review budget, sampling rate, and securing margins. Non-delegable: the decision to send a vulnerable passenger into a vehicle, alone, at night. The AI may flag it, a person must own it.
warfare, autonomous agents operating at machine speed might escalate conflicts faster than humans can comprehend or control — potentially triggering accidental catastrophes. Key Insight: Multi-agent systems create emergent harm even when every individual action passes ethical review.
Limited Capabilities Seeking Clarifications Clear boundaries on what an agent is Constraining the scope of autonomous When and how should an AI agent authorized to do, and what requires action to reduce the surface area for pause, ask, and wait, rather than proceed escalation to a human principal. unintended ethical violations. on its best inference? Most challenging from an ethical perspective.
agentic AI responsibly over the next 2–3 years? Build Supervision & Control Agents with built-in intervention points and transparency. Integrity Beyond Compliance Responsible agentic AI must be built into organizational culture.
in how responsibility, trust, and decision-making are shared between humans and machines." The Key Question for Leaders Does your organization possess the collective insight, resources, and strategic mindset necessary to lead through this transformation?
substations, meters and inverters negotiate load directly with each other Efficient, illegible Scale as attack surface No locus of accountability Agent-to-agent messages compress to codes no operator reads. Dispatch drops from minutes to seconds, and the audit trail becomes vectors, not sentences. 40,000 endpoints are 40,000 injection points. A spoofed price or tariff signal propagates fleet-wide before the first human alert is opened. A district that goes dark is the sum of thousands of local decisions, a vendor model, and a tariff API. Nobody authored it. The oversight trade: every gain in machine-to-machine efficiency removes a place where a human could have looked.
Responsibilities 2. The Authority Gradient Advisors become exception handlers. The product owner now owns pricing logic she cannot read, while the vendor owns the model behind it. Nobody can name the person who authored a given quote. The machine price becomes the default. Accepting the agent costs nothing; overriding it needs written justification and a manager. Authority quietly moves from the actuary to whoever configured the margin target. 3. The Complacency Problem 4. Automation Bias Oversight decays with success. Approval rates reach ~99% by month three, sampling falls from 10% to 2%, and the exception queue is worked at speed. SME loss-ratio drift is spotted two quarters late. Contradicting evidence gets discounted. Staff assume the model 'saw more data' and wave through a mispriced renewal, despite a claims note in the file. Complaints are answered with the model output rather than the facts. Controls that hold: a named accountable human per decision type, mandatory rotating sample audits, override and challenge rates tracked as KPIs, and a documented reason for every price a customer can contest.
Design phase and Data curation Model spec: tone, Constitutional AI: Uses a set personality, response length, of rules to promote fairness safety trade-offs and to correct injustice Validation Top concerns Position Evaluations: Safety teams ASL Levels (AI Safety Levels) Guidelines Red-teaming + Scalable Internal AI Governance Board BlenderBot & Galactica Learnings: Lessons from past failures LLaMA Models with Guardrails SAIF (Scalable Alignment Open Publishing & Open Oversight: Conducts Infrastructure Framework) Weight Releases: Supports adversarial testing misalignment at scale, unclear boundaries of model autonomous capabilities decision-making xAI as the Unsafe/ “True” transparency design Data Cards) Hallucinations, misuse, and “Neutrality” PAIR (People + AI Data Documentation (e.g., stress-test Model Cards & Usage AI Principles (since 2018) Research): UX and System Behavior Training Meta AI “Woke Bias” community auditing SAIF (Scalable Alignment Scaling safety in Infrastructure Framework) open-source and content moderation challenges Anthropic Tendency for safety The "Democratizer”. Safety can be removed