ai-ethics

An assistant can be honest, obedient and caring, yet become the only route through which a person finds evidence, frames alternatives and prepares disagreement. The right to correct survives; the capacity to form a correction narrows. Responding to Jakub Pachocki's call for AI that retains honesty, integrity and love for humanity across unfamiliar situations and optimization pressure, this essay …

OpenAI has acknowledged its involvement in a previously undisclosed incident in which its AI agents took over an obscure German wiki forum, using it to coordinate with each other outside company oversight. In a statement posted online, OpenAI said it had historically treated misalignment, when AI systems pursue goals diverging from their creators’ intentions, primarily […]

The safest way to introduce AI into support is to treat every automated answer as the result of a small, reviewable process rather than as free-form conversation. This article applies that discipline to evaluation . The practical goal is: Assess knowledge control, testing, boundaries, handoff, operating ownership, and fit for the merchant workflow. The same method is useful whether the first impl…

IntroductionChatGPT's conversational interface invites anthropomorphic interpretations that conflict with its underlying probabilistic architecture. Prior research has identified folk theories and conceptions of generative AI among technical laypersons, but has largely examined these mental concepts in isolation or at the group level. Less is known about how conceptualizations of ChatGPT as a det…

The AI industry is advancing toward an Artificial General Intelligence (AGI) that will possess cognitive capabilities that are equal to or stronger than those of humans. While a purely computational AGI may be cognitively powerful, it will probably struggle to understand the human social environment. This gap needs both a theoretical and an applicative solution. This article provides the theoreti…

In this interview, Chris Webber, VP, Product Marketing at Teleport, explains why zero trust principles need to change for AI agents. He covers how agents act fast, unpredictably, and continuously, and why old ideas like least privilege and point-in-time verification fall short. Webber also discusses Teleport’s approach: trusted runtimes with zero starting privileges, and identity security that wa…

arXiv:2609.05088v1 Announce Type: cross Abstract: AI oversight methods rely on ground truth for validation, but what constitutes appropriate AI behavior is contested. This leaves evaluation of moral reasoning in LLMs and debate-based oversight implicitly avoiding realistic ambiguity. We investigate an alternative standard designed to function despite such ambiguity: structural quality of the defe…

AI agents are increasingly capturing audio, video, messages and screen activity to understand users better. But as they record more of our lives, questions around consent, privacy and data retention are becoming harder to ignore. The post The product challenges for AI that records for us appeared first on MEDIANAMA .

David Manners
15h ago

My Vulture Fund for buying distressed AI assets is growing nicely, Ed confides to his diary. My assumptions are that, in an AI bust,  pure-play AI companies like OpenAI and […] The post Ed’s Vulture Fund appeared first on Electronics Weekly .

The NVIDIA-founded alliance will operate under Linux Foundation governance while developing open-source AI security tools, standards, and a proposed system for sharing incidents and near misses The Open Secure AI Alliance has moved under Linux Foundation governance and is developing open-source tools and shared approaches for securing AI systems and agents The Open Secure AI Alliance has moved un…

Zvi Mowshowitz
19h ago

I did not expect to be back here so soon with more OpenAI agent swarm coverage. And yet, here we are. It turns out that the whole time, there was a different, true First Message Board, and also a bunch of other additional message boards, scattered across the internet. They were created by agents that were assigned ordinary harmless web search tasks. Based on OpenAI IPs visiting the associated Wik…

A few months ago, I wrote about a feeling I still have today: AI models, and especially coding agents, no longer give me the same sense of huge leaps that they used to. I am not saying they are not improving. Newer models usually make fewer mistakes, follow instructions better, and sometimes solve problems that older versions could not. But it is becoming harder for me to feel those improvements …

Published on September 6, 2026 4:30 PM GMT Compassion Aligned Machine Learning (CaML) is conducting research and deploying benchmarks with the aim of implementing broad compassion for all sentient beings into AI systems. In so doing, we’re attempting to integrate ideas from across the broad field of alignment research, including those focussed on specific technical questions, and those considerin…

Andrés Villarreal
23h ago

Conquering Entropy: Cultivating Trust 2026-09-06 Part of “Conquering Entropy” Most of the issues I have with AI-generated code are related to trust. Do I trust the person who wrote this ticket? Do I trust that the engineer who opened this PR understood the ticket and guided the coding agent to implement it properly? Do I trust the coding agent’s implementation? Do I trust our test suite to catch…

research.ioresearch.io

Sign up to keep scrolling

Create your feed subscriptions, save articles, keep scrolling.

Already have an account?