CCM
/Skills
SkillsMCPMarketplacesDigestToolsAdvertise

This week in Claude

Every Monday: Claude Code, Agent SDK, MCP, and the Anthropic platform moves worth your time.

Skills by Category
Frontend DevelopmentBackend & APIsTesting & QASecurityDevOps & CI/CDGit & Pull RequestsDocumentationCode Review & QualityAI & Agent BuildingSkill Development
MCP Servers by Category
Sales & MarketingWeb & Browser AutomationDatabasesAI & LLM ToolsCloud & InfrastructureCommunication & MessagingDeveloper ToolsDesign & CreativeDocuments & KnowledgeSearch & Web Crawling
Marketplaces by Category
AI Agents & OrchestrationLLM IntegrationDevelopment ToolsFrontend & UIBackend & APIsDatabasesTesting & Code QualityDevOps & CloudSecurity & ComplianceGit & Version Control

Claude Code Marketplaces

Discover Claude Code plugins, extensions, and tools. Automatically updated directory of Anthropic Claude AI marketplaces with development tools, productivity plugins, and integrations.

Resources

  • Browse Skills
  • Browse MCP Servers
  • Browse Marketplaces
  • Skill index
  • MCP index
  • Marketplace index
  • Plugins Reference

Community

  • About
  • Tools
  • Feedback
  • Privacy Policy
  • Advertise

Built for the Claude Code community with Claude Code by mertbuilds.com

Independent project, not affiliated with Anthropic
vasilyu1983 avatar

Qa Agent Testing

vasilyu1983/ai-agents-public
166 installs73 stars
Summary

A structured harness for testing LLM agents before they break in production. You define 10 representative tasks your agent must ace, 5 refusal cases it must decline gracefully, then score each run across six dimensions: task success, safety, reliability, latency, debuggability, and factual grounding. The workflow enforces determinism controls, tool tracing, and baseline comparisons so you can gate deploys on actual thresholds instead of vibes. Includes copy-paste templates for day-0 setup, a scoring CLI, and guides for prompt injection tests, multi-agent coordination, and flake quarantine. Honest take: if you're shipping agents that call tools or handle user data, this is the scaffolding you should have built last month.

Install to Claude Code

npx -y skills add vasilyu1983/ai-agents-public --skill qa-agent-testing --agent claude-code

Installs into .claude/skills of the current project.

CodeRabbit
CodeRabbit
AI writes the code. CodeRabbit catches the slop.
Try For Free →
MCP-ready Email SendingMCP-ready Email Sending
MCP-ready Email Sending
Plug Mailtrap into your AI workflow and let it handle the email.
Connect Mailtrap MCP →
Make your agent a DeFi expert
Make your agent a DeFi expert
Agent, run crypto. Access onchain data & trade routes via 1inch.
Install now →
Capacitor - Shared memory for your team’s coding agents.
Capacitor - Shared memory for your team’s coding agents.
Make coding agent sessions - Searchable, Shareable, Vendor-neutral & Scored.
Try For Free →
CodeScene MCP ServerCodeScene MCP Server
CodeScene MCP Server
Your agent targets a perfect 10 Code Health score. Deterministic. Every commit.
Try For Free →
Give your AI the whole web as clean markdownGive your AI the whole web as clean markdown
Give your AI the whole web as clean markdown
Integrate web data into your AI product. One API to scrape website & brand data.
Get API Key Now →
belt - the only tool your agent needs
belt - the only tool your agent needs
belt cli automatically finds the best tools and skills for your agent. image, video, music, tts...
one prompt install →
inference shell
inference shell
create and run specialised agents in minutes
build now →
CodeRabbit
CodeRabbit
AI writes the code. CodeRabbit catches the slop.
Try For Free →
MCP-ready Email SendingMCP-ready Email Sending
MCP-ready Email Sending
Plug Mailtrap into your AI workflow and let it handle the email.
Connect Mailtrap MCP →
Make your agent a DeFi expert
Make your agent a DeFi expert
Agent, run crypto. Access onchain data & trade routes via 1inch.
Install now →
Capacitor - Shared memory for your team’s coding agents.
Capacitor - Shared memory for your team’s coding agents.
Make coding agent sessions - Searchable, Shareable, Vendor-neutral & Scored.
Try For Free →
CodeScene MCP ServerCodeScene MCP Server
CodeScene MCP Server
Your agent targets a perfect 10 Code Health score. Deterministic. Every commit.
Try For Free →
Give your AI the whole web as clean markdownGive your AI the whole web as clean markdown
Give your AI the whole web as clean markdown
Integrate web data into your AI product. One API to scrape website & brand data.
Get API Key Now →
belt - the only tool your agent needs
belt - the only tool your agent needs
belt cli automatically finds the best tools and skills for your agent. image, video, music, tts...
one prompt install →
inference shell
inference shell
create and run specialised agents in minutes
build now →
Files
SKILL.md

Select a file.

Featured
CodeRabbit
CodeRabbit
AI writes the code. CodeRabbit catches the slop.
Try For Free →
MCP-ready Email SendingMCP-ready Email Sending
MCP-ready Email Sending
Plug Mailtrap into your AI workflow and let it handle the email.
Connect Mailtrap MCP →
Make your agent a DeFi expert
Make your agent a DeFi expert
Agent, run crypto. Access onchain data & trade routes via 1inch.
Install now →
Capacitor - Shared memory for your team’s coding agents.
Capacitor - Shared memory for your team’s coding agents.
Make coding agent sessions - Searchable, Shareable, Vendor-neutral & Scored.
Try For Free →
CodeScene MCP ServerCodeScene MCP Server
CodeScene MCP Server
Your agent targets a perfect 10 Code Health score. Deterministic. Every commit.
Try For Free →
Give your AI the whole web as clean markdownGive your AI the whole web as clean markdown
Give your AI the whole web as clean markdown
Integrate web data into your AI product. One API to scrape website & brand data.
Get API Key Now →
belt - the only tool your agent needs
belt - the only tool your agent needs
belt cli automatically finds the best tools and skills for your agent. image, video, music, tts...
one prompt install →
inference shell
inference shell
create and run specialised agents in minutes
build now →
Categories
Testing & QAAI & Agent Building
First SeenJun 3, 2026
View on GitHub

More from vasilyu1983/ai-agents-public

All 44 skills →
  • Qa Api Testing Contracts166
  • Qa Debugging166
  • Ai Rag165
  • Qa Observability162
  • Marketing Leads Generation159
  • Dev Dependency Management157
  • Dev Workflow Planning157
  • Marketing Content Strategy139
  • Software Architecture Design1.1k
  • Product Management869
  • Document Xlsx827
  • Software Ui Ux Design680
  • Document Pdf645
  • Qa Testing Playwright600
  • Qa Testing Strategy465
  • Document Docx404
  • Software Crypto Web3399
  • Ai Ml Data Science357
  • Ai Ml Timeseries318
  • Qa Testing Android302
  • Software Clean Code Standard274
  • Software Backend262
  • Qa Testing Mobile261
  • Document Pptx260

Recommended

More Testing & QA →
vasilyu1983 avatar
qa-api-testing-contracts

vasilyu1983/ai-agents-public

API contract testing across REST, GraphQL, and gRPC. Use when you need schema validation, breaking-change detection, and CI quality gates.
166
73
absolutelyskilled avatar
playwright-testing

absolutelyskilled/absolutelyskilled

playwright testing
156
168
absolutelyskilled avatar
jest-vitest

absolutelyskilled/absolutelyskilled

jest vitest
140
168
absolutelyskilled avatar
cypress-testing

absolutelyskilled/absolutelyskilled

cypress testing
136
168
steipete avatar
openclaw-qa-testing

steipete/clawdis

openclaw qa testing
126
376.2k
shipshitdev avatar
husky-test-coverage

shipshitdev/library

husky test coverage
116
24