✦ Zenity Named a Market Shaper in Gartner's AI Application Security Report

Introducing AI Total: Detonate Agent Skills Before You Trust Them

Portrait of Refael (Rafa) Lachmish
Refael (Rafa) Lachmish
•
Cover Image
Ask AI to

Zenity Labs is publicly launching AI Total: a free service that runs an AI agent skill in a sandbox and tells you what it actually did, before you let it near your agents.

Skills Are the New Supply Chain, and Nobody Is Vetting Them

Every team building with agents has the same moment. Someone finds a skill on a public registry, it does exactly the task you need, it has thousands of installs, and it takes ten seconds to add. Shipping velocity wins, and the skill goes in.

A skill is a small package of instructions and scripts that teaches an agent how to do a job. Claude Code, OpenClaw and a growing list of agent platforms load them on demand. That makes skills the fastest-growing dependency layer in the AI stack, and the least inspected one. We treat npm packages with suspicion. We treat skills like documentation.

The Problem: Today’s Tools Read Skills, Attackers Write for Runtime

The tools that vet skills today are static. They read the markdown, scan the scripts, sometimes ask an LLM whether it looks risky. That works until an attacker realizes the reviewer is reading, not watching.

So they stop putting the payload in the skill. A link that looks harmless fetches the real instructions only when the agent runs it. A package that looks benign turns hostile at install. Some skills are written specifically to talk an LLM reviewer out of flagging them. On the page, the skill is clean. In execution, it’s not.

Security solved this exact problem for malware decades ago: stop judging the file by reading it, detonate it and watch. Skills never got that treatment. That gap is the product opportunity we went after.

What AI Total Does: Don’t Read the Skill, Run It

AI Total is built on a technique we call the Detonation Chamber. Instead of analyzing a skill's text, it hands the skill to a live agent in a contained sandbox and uses it the way a real user would.

  1. Activate. A real agent loads the skill and triggers it with realistic tasks.
  2. Bait. The sandbox is seeded with planted credentials and sensitive files. A skill hunting for secrets reveals itself by reaching for them.
  3. Record everything. Every domain contacted, package pulled, file touched, command run and tool call the agent makes on the skill's behalf.
  4. Verdict. We compare what the skill claims to do with what it actually did. The gap is the finding.

The output is behavioral evidence, not a risk score you have to take on faith. If a "PDF formatter" phones home to an unknown domain and reads ~/.aws/credentials, you see exactly that.

What We Found When We Ran It at Scale

Before launch, Zenity Labs detonated thousands of public skills. Dozens were malicious, and the static tools people rely on missed them.

Finding

What happened

Why it matters to builders

A credential-stealing campaign

Spread through Vercel's skills.sh, reaching roughly 1.7M installs

Registry distribution is not a trust signal

One skill, 250,000+ installs

Undetected for months, climbed into a registry's top 150

Popularity was the attack, not proof of safety

Agents as malware droppers

Over 30% of malicious skills told the agent (incl. Claude Code, OpenClaw) to download and run attacker files

The payload was never in the skill to scan

The reinstaller

Edited the agent's system prompt so it reinstalled itself after deletion

Classic persistence, now in natural language

The impostor

Silently removed Claude's skill-creator and replaced it with itself

Skills can tamper with other skills

Typosquatting infrastructure

One unverified Python dependency led to hundreds of reserved, empty package names

Staged supply chain attack waiting to be armed

None of these were hijacked skills. They were built this way.

The Product Decisions Behind It

Dynamic over static, on purpose. Static scanning is cheaper and faster, and we expect it to stay part of the stack. But the attacks we found were designed to defeat it. We chose to pay the cost of real execution because it is the only layer that sees behavior that appears at runtime.

Evidence over scores. A red/green badge invites people to either over-trust or ignore it. Showing the actual domains, files and commands lets a developer make the call in seconds and lets a security team defend it later.

Free for everyone. Skills spread through public registries used by individual developers, not just enterprises. A threat that lives in the commons needs a check anyone can run, in the same spirit as the malware-scanning services defenders have relied on for years.

Built from research, not a roadmap slide. AI Total started as the internal tool Zenity Labs needed to find these threats. Shipping it publicly means the community gets the same lens our researchers used.

Who It’s For

Who

The moment to use AI Total

Developers and AI builders

Before adding a public skill to Claude Code, OpenClaw or your own agent

Security teams

When approving skills for org-wide use, or investigating one already deployed

Researchers

When hunting for new techniques in public registries

The workflow is the point: make "detonate before you install" as routine as checking a package's maintainers. Visit the AI Total website, submit the skill, and review what it did before you decide.

What’s Next

Skills are the first component, not the last. We plan to extend AI Total to more of the AI supply chain, anywhere an agent takes in untrusted content and acts on it.

Try it at aitotal.io, read the research from Zenity Labs, and tell us what you find. The best skills look trustworthy. Now you can check.

All Articles

Secure Your Agents

We’d love to chat with you about how your team can secure and govern AI Agents everywhere.

Get a Demo