CCM
/MCP
SkillsMCPMarketplacesDigestToolsAdvertise

This week in Claude

Every Monday: Claude Code, Agent SDK, MCP, and the Anthropic platform moves worth your time.

Skills by Category
Frontend DevelopmentBackend & APIsTesting & QASecurityDevOps & CI/CDGit & Pull RequestsDocumentationCode Review & QualityAI & Agent BuildingSkill Development
MCP Servers by Category
Sales & MarketingWeb & Browser AutomationDatabasesAI & LLM ToolsCloud & InfrastructureCommunication & MessagingDeveloper ToolsDesign & CreativeDocuments & KnowledgeSearch & Web Crawling
Marketplaces by Category
AI Agents & OrchestrationLLM IntegrationDevelopment ToolsFrontend & UIBackend & APIsDatabasesTesting & Code QualityDevOps & CloudSecurity & ComplianceGit & Version Control

Claude Code Marketplaces

Discover Claude Code plugins, extensions, and tools. Automatically updated directory of Anthropic Claude AI marketplaces with development tools, productivity plugins, and integrations.

Resources

  • Browse Skills
  • Browse MCP Servers
  • Browse Marketplaces
  • Skill index
  • MCP index
  • Marketplace index
  • Plugins Reference

Community

  • About
  • Tools
  • Feedback
  • Privacy Policy
  • Advertise

Built for the Claude Code community with Claude Code by mertbuilds.com

Independent project, not affiliated with Anthropic
lfnovo avatar

Content Core

lfnovo/content-core
149
Summary

Content Core turns Claude into a universal content extraction tool by wrapping multiple engines like Firecrawl, Jina, and Crawl4AI behind a unified interface. It exposes two MCP operations: extract_content for pulling text from URLs, PDFs, videos, and audio files, and summarize_content for condensing that material with custom context. You'd reach for this when you need Claude to process external content sources during conversations, whether that's analyzing YouTube videos, summarizing research papers, or extracting insights from web articles. The server handles the complexity of choosing the right extraction engine per content type.

CodeRabbit
CodeRabbit
AI writes the code. CodeRabbit catches the slop.
Try For Free →
ego lite browserego lite browser
ego lite browser
Fastest browser for AI agents to run web automation tasks, always free.
Download Free life-time →
Granola, the best AI meeting recorder
Granola, the best AI meeting recorder
Notes, actions and memory. Without a meeting bot. First month 100% off.
Download for free →
CodeHealth MCP ServerCodeHealth MCP Server
CodeHealth MCP Server
Protect your code quality, stop the AI slop.
Try For Free →
belt - the only tool your agent needs
belt - the only tool your agent needs
belt cli automatically finds the best tools and skills for your agent. image, video, music, tts...
one prompt install →
AppSignal
AppSignal
Monitor with ease. Code with confidence.
Start Free Trial →
Agent, connect blockchain
Agent, connect blockchain
Connect your Claude agent to live crypto prices and trading routes via 1inch
Get the MCP →
Block distraction from your iPhone for freeBlock distraction from your iPhone for free
Block distraction from your iPhone for free
Block distracting apps from your iPhone permanently without a 3rd party app. Free and open source.
Block now (100% free) →
CodeRabbit
CodeRabbit
AI writes the code. CodeRabbit catches the slop.
Try For Free →
ego lite browserego lite browser
ego lite browser
Fastest browser for AI agents to run web automation tasks, always free.
Download Free life-time →
Granola, the best AI meeting recorder
Granola, the best AI meeting recorder
Notes, actions and memory. Without a meeting bot. First month 100% off.
Download for free →
CodeHealth MCP ServerCodeHealth MCP Server
CodeHealth MCP Server
Protect your code quality, stop the AI slop.
Try For Free →
belt - the only tool your agent needs
belt - the only tool your agent needs
belt cli automatically finds the best tools and skills for your agent. image, video, music, tts...
one prompt install →
AppSignal
AppSignal
Monitor with ease. Code with confidence.
Start Free Trial →
Agent, connect blockchain
Agent, connect blockchain
Connect your Claude agent to live crypto prices and trading routes via 1inch
Get the MCP →
Block distraction from your iPhone for freeBlock distraction from your iPhone for free
Block distraction from your iPhone for free
Block distracting apps from your iPhone permanently without a 3rd party app. Free and open source.
Block now (100% free) →

Content Core

License: MIT PyPI version Downloads Downloads GitHub stars GitHub forks GitHub issues Ruff

Extract, process, and summarize content from URLs, files, and text through a unified async Python API, CLI, or MCP server.

Supported Formats

CategoryFormats
WebURLs, HTML pages, YouTube videos, Reddit posts
DocumentsPDF, DOCX, PPTX, XLSX, ODT, ODS, ODP, EPUB, HTML, Markdown, plain text
MediaMP3, WAV, M4A, FLAC, OGG (audio); MP4, AVI, MOV, MKV (video)

Quick Start

pip install content-core
import content_core

result = await content_core.extract_content(url="https://example.com")
print(result.content)

Or with zero install:

uvx content-core extract "https://example.com"

CLI Usage

Content Core provides a unified content-core command with subcommands for extraction, summarization, and MCP server.

Extract

# From a URL
content-core extract "https://example.com"

# From a file
content-core extract document.pdf

# With JSON output
content-core extract document.pdf --format json

# With a specific engine
content-core extract "https://example.com" --engine firecrawl

# From stdin
echo "some text" | content-core extract

Summarize

# Summarize text
content-core summarize "Long article text here..."

# With context
content-core summarize "Long text" --context "bullet points"

# From stdin
cat article.txt | content-core summarize --context "explain to a child"

MCP Server

content-core mcp

Configuration

# Set persistent config
content-core config set llm_provider anthropic
content-core config set llm_model claude-sonnet-5

# List current config
content-core config list

# Delete a config value
content-core config delete llm_provider

Config is stored in ~/.content-core/config.toml. Priority: command flags > env vars > config file > defaults.

Zero-Install with uvx

All commands work without installation using uvx:

uvx content-core extract "https://example.com"
uvx content-core summarize "text" --context "one sentence"
uvx content-core mcp

Python API

Extraction

import content_core

# From a URL
result = await content_core.extract_content(url="https://example.com")

# From a file
result = await content_core.extract_content(file_path="document.pdf")

# From text
result = await content_core.extract_content(content="some text")

# With engine override
from content_core import ContentCoreConfig
config = ContentCoreConfig(url_engine="firecrawl")
result = await content_core.extract_content(url="https://example.com", config=config)

Summarization

import content_core

summary = await content_core.summarize("long article text", context="bullet points")

Configuration

from content_core import ContentCoreConfig

config = ContentCoreConfig(
    url_engine="firecrawl",
    document_engine="docling",
    audio_concurrency=5,
)
result = await content_core.extract_content(url="https://example.com", config=config)

MCP Integration

Content Core includes a Model Context Protocol (MCP) server for use with Claude Desktop and other MCP-compatible applications.

badge

Add to your claude_desktop_config.json:

{
  "mcpServers": {
    "content-core": {
      "command": "uvx",
      "args": ["content-core", "mcp"],
      "env": {
        "OPENAI_API_KEY": "sk-..."
      }
    }
  }
}

The MCP server exposes two tools: extract_content and summarize_content. Both return plain text.

For detailed setup, see the MCP documentation.

Agent Skill (Claude Code & Codex)

Content Core ships an Agent Skill that teaches AI agents how to use it for extracting content from external sources. This repository is also a plugin marketplace, so the skill installs natively in both harnesses.

Claude Code — add the marketplace and install the plugin:

/plugin marketplace add lfnovo/content-core
/plugin install content-core@content-core

Codex — the repository carries a Codex plugin manifest (.codex-plugin/plugin.json) and marketplace catalog (.agents/plugins/marketplace.json) pointing at the same skill.

Manual fallback — copy the skill file directly into your project:

curl -o .claude/skills/content-core/SKILL.md --create-dirs \
  https://raw.githubusercontent.com/lfnovo/content-core/main/skills/content-core/SKILL.md

Once installed, the agent can use content-core to extract content from URLs, documents, and media files — either via CLI (uvx content-core) or MCP if configured.

AI Providers

Content Core uses Esperanto to support multiple LLM and STT providers. Switch providers by changing the config — no code changes needed:

# Use Anthropic for summarization
content-core config set llm_provider anthropic
content-core config set llm_model claude-sonnet-5

# Use Groq for transcription
content-core config set stt_provider groq
content-core config set stt_model whisper-large-v3

Supported providers include OpenAI, Anthropic, Google, Groq, DeepSeek, Ollama, and more. See the Esperanto documentation for the full list.

Configuration

Content Core uses ContentCoreConfig powered by pydantic-settings. Settings are resolved in priority order: constructor args > env vars (CCORE_*) > config file (~/.content-core/config.toml) > defaults.

Environment Variables

VariableDescriptionDefault
CCORE_URL_ENGINEURL extraction engine (auto, simple, firecrawl, jina, crawl4ai)auto
CCORE_DOCUMENT_ENGINEDocument extraction engine (auto, simple, docling) — docling raises ConfigurationError if the extra is not installed; auto falls back silentlyauto
CCORE_AUDIO_CONCURRENCYConcurrent audio transcriptions (1-10)3
CCORE_AUDIO_SEGMENT_MINUTESLength of the segments long audio is split into; 0 sends the file whole10
CRAWL4AI_API_URLCrawl4AI Docker API URL (omit for local browser mode)-
CRAWL4AI_API_TOKENBearer token for the Crawl4AI Docker API (required by Crawl4AI >= 0.9.0)-
FIRECRAWL_API_URLCustom Firecrawl API URL for self-hosted instances or Firecrawl-compatible backends (e.g. fastCRW)-
CCORE_FIRECRAWL_PROXYFirecrawl proxy mode (auto, basic, stealth)auto
CCORE_FIRECRAWL_WAIT_FORWait time in ms before extraction3000
CCORE_LLM_PROVIDERLLM provider for summarization-
CCORE_LLM_MODELLLM model for summarization-
CCORE_STT_PROVIDERSpeech-to-text provider-
CCORE_STT_MODELSpeech-to-text model-
CCORE_STT_TIMEOUTSpeech-to-text timeout in seconds-
CCORE_YOUTUBE_LANGUAGESPreferred YouTube transcript languages-
CCORE_YOUTUBE_COOKIES_FILEPath to a Netscape cookies.txt for YouTube transcripts (missing file raises ConfigurationError)-
CCORE_YOUTUBE_PROXYProxy URL for YouTube transcripts, e.g. http://user:pass@host:port-

API keys for external services are set via their standard environment variables (e.g., OPENAI_API_KEY, FIRECRAWL_API_KEY, JINA_API_KEY).

Proxy Configuration

Content Core reads standard HTTP_PROXY / HTTPS_PROXY / NO_PROXY environment variables automatically. No additional configuration is needed.

YouTube on Blocked Networks

If YouTube extraction fails with IpBlocked/RequestBlocked on a flagged IP (cloud hosts, VPNs), either point CCORE_YOUTUBE_COOKIES_FILE at a cookies.txt exported from a logged-in browser (re-export it when the cookies expire), or set CCORE_YOUTUBE_PROXY to a residential proxy — datacenter proxies get blocked or CAPTCHA'd. See docs/usage.md.

Optional Dependencies

# Docling for advanced document parsing (PDF, DOCX, PPTX, XLSX)
# Required by document_engine="docling", which raises ConfigurationError without it.
# Use document_engine="auto" (default) or "simple" to proceed without Docling.
pip install content-core[docling]

# Crawl4AI for local browser-based URL extraction
pip install content-core[crawl4ai]
python -m playwright install --with-deps

# LangChain tool wrappers
pip install content-core[langchain]

# All optional features
pip install content-core[docling,crawl4ai,langchain]

Using with LangChain

When installed with the langchain extra, Content Core provides LangChain-compatible tool wrappers:

from content_core.tools import extract_content_tool, summarize_content_tool

tools = [extract_content_tool, summarize_content_tool]

Documentation

  • Usage Guide -- Python API details, configuration, and examples
  • Processors -- How content extraction works for each format
  • MCP Server -- Claude Desktop and MCP integration

Development

git clone https://github.com/lfnovo/content-core
cd content-core

uv sync --group dev

# Run tests
make test

# Lint
make ruff

License

This project is licensed under the MIT License.

Contributing

Contributions are welcome! Please see our Contributing Guide for details.

Featured
CodeRabbit
CodeRabbit
AI writes the code. CodeRabbit catches the slop.
Try For Free →
ego lite browserego lite browser
ego lite browser
Fastest browser for AI agents to run web automation tasks, always free.
Download Free life-time →
Granola, the best AI meeting recorder
Granola, the best AI meeting recorder
Notes, actions and memory. Without a meeting bot. First month 100% off.
Download for free →
CodeHealth MCP ServerCodeHealth MCP Server
CodeHealth MCP Server
Protect your code quality, stop the AI slop.
Try For Free →
belt - the only tool your agent needs
belt - the only tool your agent needs
belt cli automatically finds the best tools and skills for your agent. image, video, music, tts...
one prompt install →
AppSignal
AppSignal
Monitor with ease. Code with confidence.
Start Free Trial →
Agent, connect blockchain
Agent, connect blockchain
Connect your Claude agent to live crypto prices and trading routes via 1inch
Get the MCP →
Block distraction from your iPhone for freeBlock distraction from your iPhone for free
Block distraction from your iPhone for free
Block distracting apps from your iPhone permanently without a 3rd party app. Free and open source.
Block now (100% free) →
UpdatedFeb 7, 2026
View on GitHub