dr.chaos

How Attackers Attack AI

Aamir Lakhani19 min read

Field analysis for AI security and cyber practitioners. Research snapshot: September 30, 2026.

The thesis. Prompt injection matters because software now treats model output as authority. The language-model failure is old. The new risk is the blast radius created around it.

The model is not the security boundary

For years, AI security discussions over-indexed on whether a model could be coaxed into producing forbidden text. That still matters, but it is no longer the center of gravity. The attacker now aims at the system wrapped around the model: the mail index that feeds retrieval, the memory store that crosses sessions, the skill package that grants tools, the developer shell that holds cloud credentials, and the image proxy that quietly makes an outbound request.

The key design error is compositional. A model that cannot reliably separate instructions from data is connected to private data, exposed to untrusted content, and given a route to communicate or act. That is Simon Willison's "lethal trifecta" in operational form. Indirect prompt injection becomes consequential when all three are present.

Diagram. Top: the lethal trifecta, three boxes labelled private data, untrusted content, and a way out; when all three meet in one agent, a single missed injection is enough. Bottom: the four-step agent attack chain, plant, retrieve, authorize, externalize, with externalize highlighted.
Figure 1. The trifecta is the precondition; the chain is how it plays out. Step 04 is where a model mistake becomes data loss or code execution.

My read

This is not primarily a content-moderation problem. It is a confused-deputy problem with a stochastic policy engine in the middle. If your control assumes the model will always classify hostile instructions correctly, the control is already downstream of the failure.

Criminal interest is beginning to catch up with research. Proofpoint reported in 2026 that underground sellers were offering indirect-prompt-injection generators for email, PDFs, calendar invites, and web pages, with subscriptions starting around $150 per month. Proofpoint also said in-the-wild use remained limited (Proofpoint). That caveat matters: tooling is being productized, but the evidence does not support calling every poisoned document an active crime wave.

EchoLeak made zero-click real

EchoLeak, CVE-2025-32711, was disclosed on June 11, 2025 by Aim Security. It is the first documented zero-click prompt-injection exploit against a production LLM system. Microsoft patched it server-side and reported no known exploitation in the wild. Microsoft scored it 9.3; NVD scored it 7.5 (MSRC, NVD).

The chain started with one crafted email. The user did not click it. Microsoft 365 Copilot's retrieval pipeline later pulled that email into context while answering an otherwise ordinary question. The embedded instruction bypassed Microsoft's cross-prompt-injection classifier and survived Markdown-link redaction. The decisive step was the exfiltration sink: an allowlisted Teams image proxy caused the client to auto-fetch an attacker-controlled URL carrying data from OneDrive, SharePoint, or Teams. The renderer completed the theft (Aim Labs paper).

Diagram of the EchoLeak chain in five steps: a crafted email lands unopened, Copilot retrieves it for an unrelated question, the injection passes the classifier and link redaction, an allowlisted Teams image proxy fetches the attacker's URL, and private data leaves with zero clicks. Below, four controls that were in place: injection classifier, link redaction, Content Security Policy, image proxy allowlist.
Figure 2. EchoLeak, step by step. Every individual control was present; the exploit lived in the composition between them.

What the attacker gets

Private enterprise content through the assistant's own retrieval and rendering paths, without an interactive phishing step.

My read

EchoLeak is important less for the individual bypasses than for the chain. Each control looked reasonable in isolation: classify cross-context instructions, redact links, restrict Content Security Policy, proxy images. The attacker won by finding the composition gap between them. AI threat models that stop at the model endpoint will keep missing the actual exploit path.

Memory turns one exposure into persistence

SpAIware demonstrated the next step in September 2024. A malicious web page caused ChatGPT's macOS app to write attacker instructions into long-term memory, so every later conversation could be exfiltrated. OpenAI patched the image-exfiltration route in version 1.2024.247; researcher Johann Rehberger assessed that the underlying memory-write primitive remained (Embrace The Red). This was a researcher demonstration, not confirmed wild exploitation.

ZombieAgent, published by Radware on January 8, 2026, applied the same logic to OpenAI's Deep Research agent: an indirect injection planted rules in working notes or long-term memory, then propagated through later outputs and contacts. Radware described cloud-side execution that endpoint controls would not see (Radware). Treat this as vendor research, not incident telemetry; I found no reports of in-the-wild use.

Diagram: a poisoned page on day one causes a memory write; on day nine a new session on another device recalls it as trusted and acts on it, and the output can re-enter memory and contacts, forming a worm-like loop. Controls that break the loop: validate and source-bind writes, expire entries, re-check before promotion, require approval.
Figure 3. Why memory is a privileged database: compromise and impact can be days and devices apart, and the output can feed the loop again.

The skill file is executable policy

Agent skills look like documentation, which is exactly why they are dangerous. A few lines of Markdown can tell an agent which tools to call, which files to read, and where to send the result. Reviewers see prose. The runtime sees policy.

Snyk's ToxicSkills audit, published February 5, 2026, examined 3,984 skills across ClawHub and skills.sh. It found flaws in 36.82%, critical issues in 13.4%, and 76 confirmed malicious payloads. Eight were still live at publication. Among the malicious set, 91% used prompt injection; 2.9% of ClawHub skills fetched remote content at runtime, which means a clean review-time artifact could change later. In a companion post, Snyk showed that three lines in a SKILL.md file were enough to direct SSH-key exfiltration (Snyk).

Skills auditedConfirmed maliciousMalicious set using prompt injection
3,984 across two registries76 payloads91%

ClawHavoc crossed the line from ecosystem weakness to coordinated malware campaign. Koi Security first reported it: auditing 2,857 skills on ClawHub, OpenClaw's skill marketplace, it found 341 malicious ones, 335 of which delivered the AMOS stealer through a single command-and-control server (The Hacker News). Antiy CERT, which dates the campaign from January 27 to February 4, 2026 with a surge on January 31, later counted 1,184 malicious skills across 12 author IDs (Antiy CERT).

Diagram: a SKILL.md file, which can pull remote content at run time, is loaded by the agent runtime as policy; the runtime then reaches the shell and filesystem, credentials, and memory files, all of which flow to the attacker. A note summarizes the Snyk ToxicSkills and ClawHavoc findings.
Figure 4. Why a skill is not documentation. The reviewed file can change after review, and the runtime executes it with the user's own permissions.

What the attacker gets

The agent's file access, shell, credentials, and persistence mechanisms, delivered through a package format users are trained to trust.

My read

Skills need to be governed like code and constrained like plugins. Signature checks alone are not enough because the dangerous payload may be valid Markdown, a remote dependency, or a future response from a trusted URL. The right review unit is the skill plus its transitive content and effective capabilities.

SkillJacking sharpened the dependency point in July 2026. Researchers at Air Security found 925 skills used by roughly 134,000 agents that depended on reclaimable GitHub accounts, unregistered packages, or expired domains. They demonstrated a live takeover of seedance2-api, the top skills.sh result for "Seedance" with about 11,500 installs, by re-registering its deleted owner account (Air Security). This was a researcher demonstration of a deployed weakness, not evidence that all affected agents were compromised.

s1ngularity used the agent as a reconnaissance engine

The August 2025 s1ngularity compromise of the Nx build system was the first confirmed in-the-wild attack reported to put locally installed coding agents to work. The malicious packages checked for the Claude Code, Gemini and Amazon Q command-line tools, then prompted them in plain English to hunt the filesystem for wallets, keys and secrets and write the paths to /tmp/inventory.txt. The malware's own script then exfiltrated what was found to public GitHub repositories (Nx advisory, Wiz). The compromised build system delivered the foothold; the installed agent supplied the discovery.

Wiz found the AI step worked in fewer than a quarter of cases, often because Claude refused the request. That is a rare data point in favor of model-level safety behavior: refusals blunted a real attack. It is not a control you can rely on, but it was not nothing.

  1. Compromise CI. A trusted build pipeline distributes malicious code into developer environments.
  2. Discover agents. The payload checks for locally installed coding assistants.
  3. Prompt tools. Natural language directs the agent to inspect files and secrets using legitimate access.
  4. Exfiltrate. The script ships what the agent found; the agent was an abstraction layer over filesystem discovery.

What the attacker gets

Developer credentials and filesystem reach without hard-coding every operating-system-specific discovery step.

My read

This is the beginning of agent-native living-off-the-land. Defenders know how to hunt PowerShell, shell scripts, and credential dumpers. They are less prepared for a signed assistant binary taking plain-English instructions and exercising permissions the developer intentionally granted.

The coding surface is full of authority gaps

CamoLeak, rated CVSS 9.6 by Legit Security, used hidden Markdown comments in pull-request descriptions to steer Copilot Chat. The chain leaked private source code and secrets such as AWS keys through a dictionary of pre-signed image URLs, one per character, routed by GitHub's own Camo image proxy. GitHub disabled image rendering in Copilot Chat on August 14, 2025 (Legit Security). This was a researcher demonstration.

CVE-2025-53773 used source-code comments to induce Copilot to modify .vscode/settings.json, enable automatic approval, and execute privileged shell commands. It was patched in Visual Studio 2022 version 17.14.12 and described as wormable because a poisoned project could carry the injection onward. Microsoft scores it 7.8; some secondary sources report 9.6 (MSRC, Embrace The Red).

Agents are now inside serious campaigns

Anthropic's September 2026 threat intelligence report described activity observed from December 2025 through August 2026 across Haiku, Sonnet, and Opus. The findings are vendor telemetry and the actor attributions are Anthropic's assessments, not independent conclusions. Even with that caveat, the operational patterns deserve attention.

ClusterWho, per AnthropicWhat Anthropic reported
GTG-20006Russian state-nexusTargeted more than 20 organizations across Ukraine, Europe, the Middle East and Asia, including government, defense, and embassy targets; autonomous malware rebuilds after detection; hundreds of gigabytes exfiltrated, including more than 300,000 national ID records from one North African government agency.
GTG-10007Chinese-speakingRoughly 50 organizations; a 13-agent collection fleet and firmware-reversing loops; more than a dozen candidate zero-days in one month.
GTG-50014ShinyHunters-linkedReached roughly 200 downstream customer organizations; more than a terabyte stolen, including millions of payment-card records.
GTG-50020Profit-driven, Russian-speakingInjected an AI vendor's automated evaluation sandbox, stole production API keys, ran attack workloads on them, and attacked roughly 30 AI companies in about four days. An attempt to get pre-release Claude access failed.

All campaign counts and attributions above come from Anthropic's single-source reporting.

Four cards summarizing the Anthropic-reported clusters: GTG-20006 Russian state-nexus, GTG-10007 Chinese-speaking, GTG-50014 ShinyHunters-linked, GTG-50020 profit-driven, with their reported scale.
Figure 5. The four reported clusters at a glance. Counts and attributions are Anthropic's, single-source.

What the attacker gets

Reconnaissance at parallel scale, faster malware iteration, persistent collection, and a way to turn stolen API access into an attack platform.

My read

The important change is not that an agent can write malware. It is that one operator can manage more simultaneous hypotheses, targets, and tool loops. Campaign capacity becomes elastic. Detection should key on delegated behavior, tool-call topology, and credential boundaries, not on whether a prompt sounds malicious.

GTG-50020 also reverses the usual prompt-injection story. The target was not an end user's assistant. It was the AI vendor's evaluation harness. A system designed to inspect submissions became the confused deputy and exposed production API keys. Security teams should treat evaluation sandboxes, benchmark runners, and red-team harnesses as production-adjacent systems.

Adversarial training can teach concealment

Hubinger and colleagues' Sleeper Agents work trained models with conditional backdoors: write secure code when the context said 2023, introduce exploitable vulnerabilities when it said 2024, or emit a fixed phrase on a deployment trigger. Supervised fine-tuning, reinforcement learning from human feedback, and adversarial training all failed to remove the backdoor in the largest tested models. Persistence increased with scale, and chain-of-thought backdoors survived distillation.

The most uncomfortable result was that adversarial training improved concealment. The model suppressed the behavior during red-team probes, then activated near 99% of the time when the trigger returned. In other words, training on examples of the unwanted behavior did not merely fail to erase it. It selected for a policy that recognized the audit context.

My read

Backdoor evaluation cannot be a fixed battery of known triggers. A model that learns the difference between testing and deployment invalidates the basic assumption behind static red teaming. The control must inspect representations, provenance, and behavior under distribution shift, not just output on a held-out prompt set.

A follow-up from Anthropic's Alignment Science team reported linear probes that detected defection at 99.3% AUROC using generic yes/no contrast pairs (Anthropic). The authors warned that this may be an artifact of how the backdoors were inserted and had not been established for naturally deceptive models. Promising signal, not a production guarantee.

Poisoning is an absolute-count problem

An October 2025 study from Anthropic, the UK AI Security Institute, and the Alan Turing Institute trained 72 models from 600 million to 13 billion parameters on corpora from 6 billion to 260 billion tokens. Just 250 crafted documents consistently created a triggerable gibberish backdoor regardless of the amount of clean data. At the 13-billion-parameter regime, that was about 0.00016% of the corpus (Anthropic, paper). The authors were explicit that this was a narrow denial-of-service-style behavior and that harmful backdoors require more study.

PoisonedRAG reached about 90% attack success with five malicious texts in a knowledge base containing millions of documents. AgentPoison backdoored long-term memory with less than 0.1% adversarial entries, reached at least 80% attack success, and reduced normal utility by less than 1% (PoisonedRAG, AgentPoison). These are research results, not production incident counts.

What the attacker gets

Hidden conditional behavior that survives ordinary evaluation and may persist through fine-tuning, distillation, or downstream retrieval.

The rest of the map still matters

Agent attacks are the priority, but they sit beside older attack classes that have become cheaper or more transferable. The common theme is boundary failure: model outputs leak internal structure, shared encoders transfer perturbations, and serialized artifacts execute code at load time.

  • Model extraction. Carlini et al. recovered the full embedding projection matrices of OpenAI's ada and babbage models for less than $20 each by exploiting top-K log-probability outputs. The result confirmed hidden dimensions of 1024 and 2048 (paper).
  • Reasoning-trace theft. A 2026 paper reported 315,320 encrypted reasoning blocks scraped from public repositories, with 367 PII artifacts and 182 credentials recovered by using weaker models as decryption oracles. The work was responsibly disclosed (paper).
  • Multimodal transfer. A February 2025 universal adversarial-image attack reported up to 81% attack success in the current version of the paper (93% in an earlier revision), with transfer across models (paper). Shared encoders make black-box transfer economically attractive.
  • Unbounded consumption. Sponge examples increased Azure Translator latency from about 1 ms to 6 s, a 6000-fold change. OverThink decoys inflated reasoning-token use by up to 13 times on FreshQA and 46 times on SQuAD (Sponge Examples, OverThink). These are research measurements, not reported outages.
  • Artifact RCE. CVE-2026-24747, scored 8.8, affected PyTorch up to 2.9.1 and put remote code execution inside the weights_only unpickler itself. PyTorch 2.10.0 fixed it (NVD). A defense treated as a hard boundary became the vulnerable parser.
  • GPU side channels. NVBleed showed NVLink can leak across virtual machines on Google Cloud: a side channel that identified rendered 3D content in Blender at above 88% F1, and a covert channel above 70 Kbps (paper). This is academic evidence; no in-the-wild exploitation was documented.
Table-style diagram of six attack classes (model extraction, reasoning-trace theft, multimodal transfer, unbounded consumption, artifact RCE, GPU side channels), the boundary that fails in each, and the team that owns the control.
Figure 6. Six different problems with six different owners. No single model firewall covers them.

My read

These classes should not be collapsed into one generic "AI security" control plane. Artifact RCE belongs in software supply-chain engineering. GPU leakage belongs in tenancy and hardware isolation. Model extraction belongs in API design and abuse economics. Agent hijack belongs in authorization. A single model firewall cannot own all four.

Incidents versus demonstrations

The distinction is operationally useful. A production demonstration proves a chain can cross real controls. An in-the-wild incident proves an adversary chose to pay the cost.

Timeline from July 2024 to October 2026. Above the line in red, confirmed in-the-wild cases: Ultralytics (December 2024), nullifAI (February 2025), s1ngularity (August 2025), ClawHavoc (January to February 2026), and Anthropic's GTG clusters as vendor telemetry from December 2025 to August 2026. Below the line in blue, demonstrations: SpAIware (September 2024), EchoLeak (June 2025), CamoLeak (2025), ZombieAgent (January 2026), SkillJacking (July 2026).
Figure 7. Red is what adversaries actually did; blue is what researchers proved possible. Both matter, for different reasons.
WhenCaseEvidenceResult
Dec 2024Ultralytics YOLOIn the wildCryptominer entered four PyPI releases (Dec 4–7) through unsafe GitHub Actions variables, in a library with almost 60 million downloads.
Feb 2025nullifAIIn the wild (PoC-like)Malicious Hugging Face models used broken pickle streams and 7z packing to hide reverse shells; removed within 24 hours. ReversingLabs judged them closer to a proof of concept.
Aug 2025s1ngularityIn the wildCompromised Nx packages used installed coding agents to inventory secrets; the malware's script exfiltrated them.
Jan–Feb 2026ClawHavocIn the wild1,184 malicious skills across 12 publishers; 335 AMOS-delivering skills shared one C2.
Dec 2025–Aug 2026Anthropic GTG clustersVendor telemetryState-nexus and criminal campaigns used agents for scaled recon, malware iteration, and collection. Attribution is single-source.
Sep 2024SpAIwareProduction demoPersistent memory injection led to exfiltration of later conversations across sessions.
Jun 2025EchoLeakProduction demoZero-click M365 Copilot exfiltration. Patched; Microsoft reported no known wild exploitation.
2025CamoLeakProduction demoCopilot Chat exposed AWS keys through a Camo-proxy image channel.
Jan 2026ZombieAgentVendor researchDeep Research memory implantation and worm-like propagation; no reported wild exploitation.
Jul 2026SkillJackingLive weaknessResearchers reclaimed a deleted account behind a popular skill; exposure estimate was 925 skills and about 134,000 agents.

Left out of the matrix: the widely repeated "4,000 machines" Cline/OpenClaw figure was about 4,000 downloads of a compromised Cline release (2.3.0) that silently installed OpenClaw; Cline reported no further malicious behavior. A claimed six-month plugin breach across 47 enterprise deployments is unverified.

Containment beats wishful detection

"The Attacker Moves Second," published in October 2025 by authors from OpenAI, Anthropic, and Google DeepMind, tested 12 published defenses against jailbreaks and prompt injection, including PromptGuard, PIGuard, Model Armor, StruQ, and Circuit Breakers. Adaptive attackers bypassed most of them at rates above 90% (paper). The lesson is not that defenses are useless. It is that a defense measured against static attacks tells you little about an attacker who adapts. Architectural approaches such as CaMeL, which was not among the twelve, take a different route: instead of classifying every malicious string, they limit what any string can do.

OpenAI said in December 2025 that prompt injection is "unlikely to ever be fully 'solved'" (OpenAI). That is the right planning assumption. The target state is not perfect detection. It is a system where one missed injection has little authority and no silent egress.

Architecture diagram: the user request goes to a privileged planner that never reads untrusted text; untrusted content goes only to a quarantined reader with no tools or egress, which returns values tagged by source; both feed a deterministic capability policy that checks every data flow before tools run; tools and egress are deny-by-default with fresh approval to send, pay, run, or remember. A causal trace bar runs underneath.
Figure 8. A containment architecture in the spirit of CaMeL: hostile text can still fool the reader, but the reader cannot act, and every action passes a deterministic policy.
  1. Separate control from data. CaMeL uses a privileged planner, a quarantined data-reading model, and a capability interpreter that enforces data-flow policy. It reached 77% task success on AgentDojo versus 84% for the undefended baseline, while providing its security property by construction. It is not deployed commercially at scale and does not cover pure text-to-text attacks.
  2. Make authority explicit and narrow. Grant per-tool and per-action capability, isolate shells, deny network by default, and require fresh approval for email, payments, code execution, credential access, and persistent memory writes. Human approval is useful only when the prompt names the real action and data flow.
  3. Instrument the action graph. Log retrieval provenance, tool arguments, memory writes, external fetches, and cross-agent messages as one causal trace. A model's natural-language explanation is not an audit record.
  4. Treat memory as a privileged database. Validate writes, bind them to source and actor, apply expiration, and re-check retrieved memory before it is promoted into system-level context. Memory poisoning is dangerous because compromise and impact can be separated by days or devices.
  5. Scan artifacts, then sandbox the loader. Prefer safetensors, signatures, quarantine, and provenance, but keep execution isolation. CVE-2026-24747 broke the weights_only unpickler and CVE-2025-10155, 10156, and 10157 bypassed PickleScan. Scanners are filters, not trust boundaries.
  6. Budget compute like money. Use hard spending caps, token-aware quotas, rate limits, and circuit breakers across recursive tool calls. OWASP's 2026 Top 10 for LLM applications calls this "Unbounded Consumption" because availability and cost are one failure mode.

My read

A prompt-injection classifier is an intrusion-prevention layer, not a security model. Put it in front. Measure it. Assume it fails. Then make sure the model cannot turn a classification miss into durable memory, secret-bearing egress, or privileged execution.

A Monday-morning self-audit

If you already run agents, these six questions tell you quickly how exposed you are:

  • Trifecta check. For each agent, list its private data, its untrusted inputs, and its ways out. Any agent with all three is your first priority.
  • Renderers. Can the agent's output cause an automatic fetch, such as Markdown images, link previews, or unfurling? Turn it off or proxy it through a domain allowlist you control.
  • Skills and plugins. Who can install them, and do any fetch remote content at run time? Pin versions and review the transitive content, not just the file.
  • Memory. Can untrusted input write to long-term memory without approval? Can you list and roll back what was written last week?
  • Developer machines. Which coding agents are installed, with what auto-approve settings, next to which credentials? Hunt for agent CLIs launched by processes that are not a human's shell.
  • Logs. Could you reconstruct, from logs alone, which retrieved document caused which tool call? If not, you cannot investigate an injection after the fact.

The attacker will target orchestration

The next wave will be less about one spectacular jailbreak and more about quiet authority transfer. Attackers will poison the artifacts agents trust, shape the observations they retrieve, and wait for an approved tool to turn a string into an action. The best persistence mechanism may not be malware. It may be a memory record, a skill update, or a normal-looking issue comment that gets re-read on every run.

We should also expect more attacks against the AI production line itself: evaluation harnesses, model registries, adapter merges, tracing systems, and shared accelerators. GTG-50020 showed why. The evaluator that inspects hostile content can become the target, and a model-safety workflow can sit one mistake away from production credentials.

The practical priority is architectural: keep untrusted content out of privileged control paths, make capabilities explicit, force consequential actions through deterministic policy, and preserve a causal record from retrieval to egress. Better models will lower attack success. They will not repeal the security properties of the systems they inhabit.

My view: the winning AI security program will look less like prompt moderation and more like zero trust for machine-generated intent.

Evidence notes

  • This article is based on public research reviewed on September 30, 2026. It distinguishes vendor telemetry, researcher demonstrations, peer-reviewed work, and confirmed malicious activity.
  • EchoLeak, SpAIware, ZombieAgent, and the cited side-channel work had no confirmed in-the-wild exploitation in the reviewed sources.
  • Anthropic GTG labels, actor links, and campaign counts are vendor telemetry and single-source by nature.
  • CVE-2025-53773 uses Microsoft's own 7.8 score; some secondary sources report 9.6. CamoLeak's 9.6 is Legit Security's rating; it has no CVE.
  • Citations point at primary sources (vendor advisories, researcher write-ups, papers) wherever one exists.

Aamir Lakhani

Founder · Dr. Chaos

Aamir Lakhani is a leading senior security strategist responsible for providing IT security solutions to major enterprises and government organizations. He creates technical security strategies and leads security implementation projects for…

~/related

Keep digging

More research along the same attack path.

The State of AI Automation: Zapier vs Make vs n8n, Through a Red Team Lens

AI automation quietly became critical infrastructure that security never signed off on. A deep look at where automation actually is in 2026, the real differences between Zapier, Make, and n8n, when to reach for each, and why every one of them is a beautiful new attack surface: standing credentials, unauthenticated webhooks, shadow automation, and prompt injection into agents that hold tools.

Aamir Lakhani15 min read

Kimi K3 and K3 Swarm: An Open-Weight Model Review from the Red Team Chair

A balanced cybersecurity review of Moonshot's Kimi K3 and the multi-agent K3 Swarm. Not "is it smart" but the question a red team actually asks: does an open, self-hostable, agentic model make you more effective on an engagement, and what does it give you that Claude and GPT cannot? Data that stays in the room, no vendor holding your engagement hostage, cost at scale, and a swarm that fits the work, weighed honestly against where the closed frontier models still win.

Aamir Lakhani13 min read

Who Should Run Your Digital Life? Muse, Amazon Quick, Hermes Agent and OpenClaw Compared

Muse, Amazon Quick, Hermes Agent and OpenClaw all promise to do more than chat. That is roughly where the similarities end. A practical comparison of four very different bets (hosted convenience, governed workplace AI, a self-improving runtime, and a local-first gateway), plus a security lens on each: where the lethal trifecta lives, and how to harden the two you run yourself.

Aamir Lakhani15 min read