News · 2026-10-10
A new study finds agent-skill copies rarely receive fixes and stars miss distribution hubs
Researchers behind Skill Constellations reconstructed how agent skills spread between GitHub repositories and found that copied skills seldom receive later fixes, while repository stars poorly identify the sources that spread them. In a manually labeled sample of 200 skills, the study’s strict risk flag achieved 98.3 percent precision and 80.9 percent recall. The findings make skill installation a software-supply-chain question, without showing that flagged skills are necessarily malicious.
Key facts
- The authors report 98.3 percent precision and 80.9 percent recall for a strict flag on 200 manually labeled skills.
- The paper’s first version is dated October 8, 2026.
- It covers public GitHub history from the format’s first ten months and excludes forks.
- The primary source is Skill Constellations, with full text and author code and data.
An agent skill can look like a helpful set of instructions. It can also contain scripts, tool permissions, and executable behavior that runs with the user’s access. Installing a collection of skills therefore resembles installing software more than bookmarking advice. The host’s resources and the way the agent follows those instructions determine what the installation can affect.
Fahd Seddik’s linked project and the paper name the subject “Tracing the Supply Chain of Agent Skills on GitHub.” That framing is the contribution: the study looks at movement and updates, not just a pile of files counted at one moment. It asks which repository copied a skill, which earlier repository appears to have supplied it, and when the transfer occurred.
The researchers reconstruct the network from Git histories and identify copying relationships through a threshold based on shared lineages. They check part of the date and source reconstruction against installer records. The inferred source is a distribution source, not proof of original authorship. A repository can be influential because it republishes other people’s work, even if it created none of the underlying instructions.
Think of a building where everyone photocopies an emergency procedure from a shared binder. If the original procedure is corrected, the photocopies do not automatically update. The most popular binder is not necessarily the one that supplied the most copies. The risk lies in the copying network and the missing correction path. The paper finds analogous problems in agent-skill distribution.
That is why stars are an imperfect audit priority. Stars measure a kind of public attention, while the copy network measures how content travels. A smaller repository can sit upstream of many later adoptions. A large, popular repository can mostly collect content from elsewhere. Auditing the network’s distribution hubs can therefore reach a different set of future installations than auditing the repositories with the largest star totals.
The study tests that idea on a held-out later period. It estimates how many subsequent adoptions could be prevented if flagged skills were removed from repositories selected for audit. This is a simulation of an intervention, not evidence that a live audit stopped attacks. New skills outside the known graph can still account for a substantial share of later risky adoption.
The risk labels also need careful interpretation. They flag capabilities such as bundled executables, pre-approved tools, and risky instructions. Those capabilities may have legitimate uses. A useful automation script and a harmful script can both execute commands. The label indicates where closer review is warranted; it does not demonstrate malware or a deliberate attacker.
The reported precision and recall illustrate two different limitations. High precision means the strict flag was usually correct when it flagged the labeled capability in the sample. The lower recall means it missed some labeled cases. The authors describe the risk counts as lower bounds. An unflagged repository should not inherit a clean security verdict from that result.
The implications apply to shipping ecosystems such as Anthropic’s knowledge-work plugins, which combine skills with connectors, commands, and other workflow components. The paper does not accuse that collection of wrongdoing. It supplies a reason to ask any ecosystem about provenance, version references, update propagation, and the permissions an installed component receives.
Ground Truth’s data-poisoning lesson explains a related trust problem at the training-data boundary. Skills create another boundary at runtime, where copied instructions can shape tool use without changing model weights. The agent-identity lesson explains why reducing granted authority can limit the consequences of a bad component.
The strongest caveat is the study’s scope: an inferred public-GitHub network, a particular risk-labeling approach, and a simulated later-period audit. It does not measure every private installation or establish the prevalence of malicious skills. Its actionable finding is narrower and useful: copy-based installation needs a reliable way to identify origin and receive fixes, and popularity alone is a weak substitute for that information.
Key questions
Does Skill Constellations show that flagged agent skills are malicious?
How accurate was the study’s strict risk flag?
Did the proposed audits actually stop harmful installations?
Cite this
APA
Ground Truth. (2026, October 10). A new study finds agent-skill copies rarely receive fixes and stars miss distribution hubs. Ground Truth. https://groundtruth.day/news/skill-constellations-agent-skills-copy-supply-chain.html
BibTeX
@misc{groundtruth:skill-constellations-agent-skills-copy-supply-chain,
title = {A new study finds agent-skill copies rarely receive fixes and stars miss distribution hubs},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/news/skill-constellations-agent-skills-copy-supply-chain.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.