Dev News Daily ENDE

The same week an AI agent found 24 Android bugs, Red Hat asked who fences the agent

Two posts published on 28 September describe the same technology from opposite ends: an AI agent as a tool that finds security flaws, and an AI agent as a system that has to be contained.

The agent as auditor. GitHub's Security Lab says it has reported more than 20 vulnerabilities in Android applications using its open-source Taskflow Agent, a way of packaging prompts and multi-step workflows so that a model works through an audit in stages. The researcher added a taskflow that separates mobile entry points, the places attacker-controlled data can arrive, from everything else in a repository, and another that makes the model check each entry point against a fixed list of Android vulnerability classes. The reason given is telling: mobile vulnerabilities are less widely known, and model output is not deterministic, so the essential checks are written down rather than left to the model. A run takes an hour or two on a medium-sized repository and needs a Copilot licence with premium model requests.

The agent as risk. In a post on the same day, Red Hat's chief technology officer argues that once agents can call APIs, use credentials and change business systems, model safeguards cannot carry the security burden alone. The controls he lists sit outside the model: the agent's identity and permissions, its access to tools and data, its runtime, its network connections. They should, he writes, assume the agent may behave unexpectedly, and limit what it can reach even then.

The same week an AI agent found 24 Android bugs, Red Hat asked who fences the agent
The same week an AI agent found 24 Android bugs, Red Hat asked who fences the agent — Dev News Daily

What connects them

Both posts reach the same design rule by different routes. GitHub's audits work because a human fixes the steps and the checklist instead of trusting the model to remember them. Red Hat's argument is that permissions must be enforced by the surrounding system instead of trusting the model to respect them. In both cases the model is useful, and in both cases the reliable part is the structure placed around it. For teams deploying agents, that suggests a simple test: list which of the agent's safeguards would still hold if the model ignored its instructions. The ones that would are the controls; the rest are hopes.

Written by Victoria Shinder.