有些任务需要另一个 Agent,有些需要用户作出决定,还有些只是等待外部事件。本实验把这些协作方式作为工具,学习主 Agent 怎样交出任务并接收结果。
建议按以下顺序阅读:理解问题与方法 → 准备环境与输入 → 按照步骤完成实验 → 分析结果与形成判断 → 阅读实现与继续探索 → 排查问题与查阅资料。
协作调用不仅传递一个问题,还需要明确上下文、预期结果和完成条件。过少的上下文使协作者无法判断,过多的上下文又增加成本和干扰。等待用户回复与等待子任务完成也应有不同的状态。
为 AI Agent 提供协作能力的综合 Model Context Protocol(MCP)服务器,涵盖浏览器自动化、人机协同、通知与定时器管理。
- 导航 URL、管理标签页
- 抽取网页内容
- 用 AI Agent 执行高层浏览器任务
- 截图
- 完整虚拟浏览器能力
- 以 sync(等待结果)或 async(返回
task_id)模式 spawn 子 Agent - 向子 Agent 发送后续消息、取消运行中的子 Agent
- 两种上下文传递策略(可检查上下文文本与 token 数):
minimal— 只传任务 + 可选手选片段(最省、隐私好,可能饿死子 Agent)llm_generated— 额外一次 LLM 调用,从父轨迹合成紧凑、隐私过滤的交接上下文
- 子 Agent system prompt 使用带标签的上下文来源(
[FROM_MAIN_AGENT]/[FROM_USER]/[TOOL_RESULT])与标准化 JSON 输出
- 敏感操作请求管理员审批
- 向人类管理员请求输入
- 管理待处理审批
- 可配置超时与通知渠道
- 经 SMTP 或 SendGrid 发信
- 支持 HTML
- 抄送与附件
- 灵活配置
- Telegram bot
- Slack webhook
- Discord webhook
- 可配置默认频道
- 一次性定时器
- 循环定时器
- 取消与管理
- 持久化存储
- 到期回调通知
mcp_browser_navigate— 导航到 URLmcp_browser_get_content— 获取页面内容mcp_browser_execute_task— 执行 AI 驱动的浏览器任务mcp_browser_screenshot— 截图mcp_browser_list_tabs— 列出标签页
mcp_send_email— 发送邮件mcp_send_telegram_message— Telegram 消息mcp_send_slack_message— Slack 消息mcp_send_discord_message— Discord 消息
mcp_spawn_subagent— 创建子 Agent(sync/async,minimal/llm_generated上下文)mcp_send_message_to_subagent— 向子 Agent 发后续消息mcp_cancel_subagent— 取消子 Agentmcp_get_subagent_status— 查询状态/结果(async)
mcp_request_admin_approval— 请求管理员审批mcp_request_admin_input— 请求管理员输入mcp_respond_to_request— 响应审批请求(管理员侧)mcp_list_pending_requests— 列出待处理请求
mcp_set_timer— 一次性定时器mcp_set_recurring_timer— 循环定时器mcp_cancel_timer— 取消定时器mcp_list_timers— 列出定时器mcp_get_timer_status— 查询定时器状态
服务器按模块组织:
collaboration-tools/
├── src/
│ ├── main.py # MCP server entry point
│ ├── config.py # Configuration management
│ ├── browser_tools.py # Browser automation
│ ├── notification_tools.py # Email & IM notifications
│ ├── hitl_tools.py # Human-in-the-loop
│ └── timer_tools.py # Timer management
├── requirements.txt # Python dependencies
├── env.example # Example configuration
└── README.md # This file
先从本地示例开始。依赖安装可能需要联网,但下面标明的离线路径不需要模型 API Key。若随后切换到真实模型,请再完成相应的服务配置。
- 从仓库根目录安装并激活统一的第 4 章环境:
# 在仓库根目录使用统一的第 4 章环境
uv sync --locked --python 3.12 --extra ch4
# 切换目录前先激活环境:
# macOS/Linux:
source .venv/bin/activate
# Windows PowerShell:.venv\Scripts\Activate.ps1
# Windows cmd:.venv\Scripts\activate.bat
# 未安装 uv 时可用 pip 兜底:
# python -m pip install -e ".[ch4]"
cd chapter4/collaboration-tools
# 精确复现旧版单项目环境,含直接 Playwright/pydantic-settings/scheduler 约束:
# python -m pip install -r requirements.txt- 复制环境模板并配置:
cp env.example .env
# Edit .env with your configuration- 安装 Playwright 浏览器(浏览器自动化):
playwright install chromium在 .env 中设置环境变量:
BROWSER_HEADLESS=false
BROWSER_USER_DATA_DIR=~/.config/collaboration-tools/browser# SMTP (Gmail example)
SMTP_HOST=smtp.gmail.com
SMTP_PORT=587
SMTP_USERNAME=your-email@gmail.com
SMTP_PASSWORD=your-app-password
SMTP_FROM_EMAIL=your-email@gmail.com
# Or use SendGrid
SENDGRID_API_KEY=your-sendgrid-api-keyTELEGRAM_BOT_TOKEN=your-telegram-bot-token
TELEGRAM_DEFAULT_CHAT_ID=your-chat-id
SLACK_WEBHOOK_URL=https://hooks.slack.com/services/YOUR/WEBHOOK
DISCORD_WEBHOOK_URL=https://discord.com/api/webhooks/YOUR/WEBHOOKHITL_ADMIN_EMAIL=admin@example.com
HITL_TIMEOUT_SECONDS=3600OPENAI_API_KEY=your-openai-api-key
OPENAI_MODEL=gpt-5.6-lunaOpenRouter 通用兜底:所有 LLM 入口(
spawn_subagent、 intelligence 工具、browser-use)经src/llm_fallback.py解析凭据。 未设置OPENAI_API_KEY但设置了OPENROUTER_API_KEY时,走 OpenRouter(base_url=https://openrouter.ai/api/v1,模型 id 映射为provider/model,如gpt-5.6-luna→openai/gpt-5.6-luna)。两者皆无时, 子 Agent 以确定性离线模式运行(不编造输出)。
- 根目录
ch4安装使用 Python 3.12(browser-use要求 Python 3.11+) - OpenAI API key(浏览器 AI 任务)
- 可选:邮件/IM 凭据
- Playwright 浏览器(浏览器自动化)
先列出工具目录,按用途区分子任务、人机协作和通知。再阅读下文中的本地演示,跟踪主任务何时等待、拿到什么结果、怎样恢复执行。通知类集成只在配置了明确接收对象后使用。
不启动 MCP 服务器,也可以用统一的命令行入口列出、单独调用协作工具,或运行端到端演示。
帮助信息为中文,-h 可查看任意子命令的参数:
python main.py --help # 总览
python main.py list # 列出全部协作工具(子 Agent / HITL / 多渠道通知)
python main.py demo # 端到端协作演示:客服协调 Agent 处理一笔退款
python main.py subagent -h # 子 Agent 子命令帮助
python main.py hitl -h # HITL 子命令帮助
python main.py notify -h # 通知子命令帮助常用示例:
# 对比两种上下文传递策略(minimal vs llm_generated)
python main.py subagent compare
# 创建子 Agent(同步、最小化上下文)
python main.py subagent spawn --task "查询订单 A12345 状态" --strategy minimal --role 订单查询助手
# 关键决策请求管理员批准;--auto-approve 在后台模拟管理员应答,便于离线演示闭环
python main.py hitl approve --message "删除 1000 条记录?" --timeout 5 --auto-approve
# 多渠道通知
python main.py notify slack --message "部署完成"demo 会串联三类协作工具:① 委派子 Agent 审批退款并对比上下文策略;② 大额操作
触发 HITL 审批(演示"超时前批准"与"超时保守默认"两种路径);③ 向协作者多渠道通知结果。
其中 HITL 与通知路径完全离线可跑;子 Agent 的真实执行与 llm_generated 策略需要
OPENAI_API_KEY(未配置时会明确提示,命令仍可正常解析运行)。
使用 stdio 传输启动:
python src/main.py也可作为 MCP 服务器接入任意兼容客户端。
python quickstart.py对同一任务分别用两种上下文传递策略 spawn,并打印差异(交接 token、额外准备成本、
是否泄漏隐私字段、各子 Agent 结果)。需要 OPENAI_API_KEY
(默认模型 gpt-5.6-luna,可用 OPENAI_MODEL 覆盖):
export OPENAI_API_KEY=your-openai-api-key
python subagent_comparison.py通常 minimal token 更少且不泄漏隐私字段,但子 Agent 可能返回 need_info;
llm_generated 多一次 LLM 调用交接更丰富、经隐私过滤的上下文,便于子 Agent 完成任务。
在 Claude Desktop 配置(claude_desktop_config.json)中加入:
{
"mcpServers": {
"collaboration-tools": {
"command": "python",
"args": ["/path/to/collaboration-tools/src/main.py"],
"env": {
"OPENAI_API_KEY": "your-key-here"
}
}
}
}# Navigate to a website
await mcp_browser_navigate(url="https://example.com")
# Execute a complex task
await mcp_browser_execute_task(
task="Search for AI agent tutorials on Google and extract the top 5 results"
)
# Take a screenshot
await mcp_browser_screenshot(full_page=True)# Send email
await mcp_send_email(
to_email="user@example.com",
subject="Task Completed",
body="Your task has finished successfully!"
)
# Send Slack message
await mcp_send_slack_message(
message="🎉 Deployment successful!"
)# Request approval for sensitive action
result = await mcp_request_admin_approval(
request_message="Delete 1000 records from database?",
urgent=True,
timeout_seconds=300
)
if result["approved"]:
# Proceed with action
pass# Set a timer
await mcp_set_timer(
duration_seconds=300,
timer_name="Check website",
callback_message="Time to check the website status"
)
# Set recurring timer
await mcp_set_recurring_timer(
interval_seconds=3600,
max_occurrences=24,
timer_name="Hourly health check"
)检查结果是否足以让主任务继续,而不只看协作者是否返回了一段文字。对于需要用户决定的事项,应能辨认决定是否真正收到,不能把等待时间当作同意。
把任务交给子 Agent 时,哪些背景必须提供,哪些信息可以让它自行查询?
Companion code for AI Agents in Depth, Chapter 4 — Experiment 4-5 ★★. MCP server: browser automation, sub-agents, HITL, multi-channel notifications, timers.
配套《深入理解 AI Agent》第 4 章 实验 4-5 ★★。协作 MCP 服务器:浏览器、子 Agent、HITL、多渠道通知、定时器。
若浏览器自动化失败:
# Reinstall Playwright browsers
playwright install chromium --force- Gmail 请使用 应用专用密码
- 不要开启「不够安全的应用访问」(改用应用专用密码)
- 通过 @BotFather 创建 bot
- 用 @userinfobot 获取 chat ID
若出现 "ChatOpenAI is not fully defined" 或 Pydantic 校验错误:
- 这是 LangChain 与 Pydantic v2 的已知兼容问题
- 修复:ChatOpenAI 仅在需要时按需初始化(
browser_execute_task) - 简单导航不需要 OpenAI API key
- 仅自主浏览器任务(
browser_execute_task)需要OPENAI_API_KEY
MIT License
欢迎提交 issue 或 pull request。
- HITL + notify paths in
python main.py demorun offline without API keys. python main.py demo中 HITL 与通知路径可离线、无需 API Key。- Browser AI tasks and
llm_generatedsub-agent strategy need an LLM key. - 浏览器 AI 任务与
llm_generated子 Agent 策略需要 LLM Key。
A comprehensive Model Context Protocol (MCP) server that provides collaboration tools for AI agents, including browser automation, human-in-the-loop assistance, notifications, and timer management.
- Navigate to URLs and manage browser tabs
- Extract content from web pages
- Execute high-level browser tasks using AI agents
- Take screenshots
- Full virtual browser capabilities
- Spawn sub-agents in sync (wait for result) or async (returns a
task_id) mode - Send follow-up messages to a sub-agent and cancel a running one
- Two context-passing strategies, made inspectable (context text + token count):
minimal— pass only the task plus an optional hand-picked slice (cheapest, private, may starve the sub-agent)llm_generated— one extra LLM call synthesizes a compact, privacy-filtered hand-off context from the parent trajectory
- Sub-agent system prompt uses labeled context sources (
[FROM_MAIN_AGENT]/[FROM_USER]/[TOOL_RESULT]) and standardized JSON output
- Request admin approval for sensitive actions
- Request input from human administrators
- Manage pending approval requests
- Configurable timeout and notification channels
- Send emails via SMTP or SendGrid
- Support for HTML emails
- CC recipients and attachments
- Flexible configuration
- Telegram bot integration
- Slack webhook support
- Discord webhook support
- Configurable default channels
- Set one-time timers
- Create recurring timers
- Cancel and manage timers
- Persistent timer storage
- Callback notifications when timers expire
- Install and activate the shared Chapter 4 environment from the repository root:
# From the repository root: use the shared Chapter 4 environment
uv sync --locked --python 3.12 --extra ch4
# Activate it before changing directories:
# macOS/Linux:
source .venv/bin/activate
# Windows PowerShell: .venv\Scripts\Activate.ps1
# Windows cmd: .venv\Scripts\activate.bat
# pip fallback when uv is not installed:
# python -m pip install -e ".[ch4]"
cd chapter4/collaboration-tools
# Exact legacy parity path, including direct Playwright/pydantic-settings/scheduler pins:
# python -m pip install -r requirements.txt- Copy the example environment file and configure it:
cp env.example .env
# Edit .env with your configuration- Install Playwright browsers (for browser automation):
playwright install chromiumConfigure the server by setting environment variables in .env:
BROWSER_HEADLESS=false
BROWSER_USER_DATA_DIR=~/.config/collaboration-tools/browser# SMTP (Gmail example)
SMTP_HOST=smtp.gmail.com
SMTP_PORT=587
SMTP_USERNAME=your-email@gmail.com
SMTP_PASSWORD=your-app-password
SMTP_FROM_EMAIL=your-email@gmail.com
# Or use SendGrid
SENDGRID_API_KEY=your-sendgrid-api-keyTELEGRAM_BOT_TOKEN=your-telegram-bot-token
TELEGRAM_DEFAULT_CHAT_ID=your-chat-id
SLACK_WEBHOOK_URL=https://hooks.slack.com/services/YOUR/WEBHOOK
DISCORD_WEBHOOK_URL=https://discord.com/api/webhooks/YOUR/WEBHOOKHITL_ADMIN_EMAIL=admin@example.com
HITL_TIMEOUT_SECONDS=3600OPENAI_API_KEY=your-openai-api-key
# Or use Alibaba Cloud Model Studio / Bailian (Qwen):
# COLLAB_PROVIDER=dashscope # qwen and bailian are aliases
# DASHSCOPE_API_KEY=your-dashscope-api-key
OPENAI_MODEL=gpt-5.6-lunaUniversal OpenRouter fallback: all LLM entry points (
spawn_subagent, intelligence tools, browser-use) resolve credentials viasrc/llm_fallback.py. WhenOPENAI_API_KEYis absent butOPENROUTER_API_KEYis set, they route through OpenRouter (base_url=https://openrouter.ai/api/v1, model id mapped toprovider/modelform, e.g.gpt-5.6-luna→openai/gpt-5.6-luna). With neither key set, sub-agents run in deterministic offline mode (no fabricated output).
Without starting the MCP server, use the unified CLI to list tools, call them individually, or run end-to-end demos. Help text is Chinese; -h works on any subcommand:
python main.py --help # overview
python main.py list # list all collaboration tools (sub-agent / HITL / multi-channel notify)
python main.py demo # end-to-end collab demo: support agent handles a refund
python main.py subagent -h # sub-agent subcommand help
python main.py hitl -h # HITL subcommand help
python main.py notify -h # notify subcommand helpCommon examples:
# Compare two context-passing strategies (minimal vs llm_generated)
python main.py subagent compare
# Spawn sub-agent (sync, minimal context)
python main.py subagent spawn --task "查询订单 A12345 状态" --strategy minimal --role 订单查询助手
# Sensitive decision needs admin approval; --auto-approve simulates admin reply offline
python main.py hitl approve --message "删除 1000 条记录?" --timeout 5 --auto-approve
# Multi-channel notification
python main.py notify slack --message "部署完成"The formal Experiment 4-5 runner defaults to credential-free notification
preflights. Use --interactive-human to pause on a real pending MCP approval
and accept exactly one live APPROVE or REJECT line from standard input. Use
--real-notifications only when email, Telegram, and Slack are all configured;
the runner fails before creating a run directory if any channel is missing and
redacts credentials and delivery identifiers from retained receipts. The
context comparison deliberately retains a hard-coded, non-secret privacy canary
in its input receipt so the validator can prove that it is absent from both
prepared handoffs. publication_authorized records only whether MCP accepted a
live approval to publish that run's validation artifact; it does not imply that
the experiment passed or that official_complete is true.
python run_experiment_4_5.py \
--campaign-id real_mcp_human_example \
--interactive-human \
--human-timeout-seconds 14400
python validate_experiment_4_5.py \
validation/experiment_4_5/real_mcp_human_exampledemo chains three collaboration tool types: (1) delegate a sub-agent for refund approval and compare context strategies; (2) large action triggers HITL (approve-before-timeout vs conservative default-on-timeout); (3) multi-channel notify collaborators. HITL and notify paths run fully offline; real sub-agent execution and llm_generated need OPENAI_API_KEY (if unset, the command still parses and runs with a clear prompt).
Start the server using stdio transport:
python src/main.pyOr use it as an MCP server with any MCP-compatible client.
Run the quickstart demo to see all features in action:
python quickstart.pySpawn a sub-agent under both context-passing strategies on the same task and
print the difference (context tokens handed off, extra preparation cost, whether
private data leaked, and each sub-agent's result). Requires OPENAI_API_KEY
(default model gpt-5.6-luna, override with OPENAI_MODEL):
export OPENAI_API_KEY=your-openai-api-key
python subagent_comparison.pyTypically minimal uses far fewer tokens and never leaks private fields, but the
sub-agent may return need_info; llm_generated spends one extra LLM call to
hand off richer, privacy-filtered context so the sub-agent can complete the task.
Add to your Claude Desktop configuration (claude_desktop_config.json):
{
"mcpServers": {
"collaboration-tools": {
"command": "python",
"args": ["/path/to/collaboration-tools/src/main.py"],
"env": {
"OPENAI_API_KEY": "your-key-here"
}
}
}
}mcp_browser_navigate- Navigate to a URLmcp_browser_get_content- Get page contentmcp_browser_execute_task- Execute AI-driven browser taskmcp_browser_screenshot- Take a screenshotmcp_browser_list_tabs- List all open tabs
mcp_send_email- Send email notificationmcp_send_telegram_message- Send Telegram messagemcp_send_slack_message- Send Slack messagemcp_send_discord_message- Send Discord message
mcp_spawn_subagent- Spawn a sub-agent (sync/async,minimal/llm_generatedcontext)mcp_send_message_to_subagent- Send a follow-up message to a sub-agentmcp_cancel_subagent- Cancel a sub-agentmcp_get_subagent_status- Get a sub-agent's status/result (for async)
mcp_request_admin_approval- Request admin approvalmcp_request_admin_input- Request admin inputmcp_respond_to_request- Respond to approval request (admin)mcp_list_pending_requests- List pending requests
mcp_set_timer- Set a one-time timermcp_set_recurring_timer- Set a recurring timermcp_cancel_timer- Cancel a timermcp_list_timers- List all timersmcp_get_timer_status- Get timer status
# Navigate to a website
await mcp_browser_navigate(url="https://example.com")
# Execute a complex task
await mcp_browser_execute_task(
task="Search for AI agent tutorials on Google and extract the top 5 results"
)
# Take a screenshot
await mcp_browser_screenshot(full_page=True)# Send email
await mcp_send_email(
to_email="user@example.com",
subject="Task Completed",
body="Your task has finished successfully!"
)
# Send Slack message
await mcp_send_slack_message(
message="🎉 Deployment successful!"
)# Request approval for sensitive action
result = await mcp_request_admin_approval(
request_message="Delete 1000 records from database?",
urgent=True,
timeout_seconds=300
)
if result["approved"]:
# Proceed with action
pass# Set a timer
await mcp_set_timer(
duration_seconds=300,
timer_name="Check website",
callback_message="Time to check the website status"
)
# Set recurring timer
await mcp_set_recurring_timer(
interval_seconds=3600,
max_occurrences=24,
timer_name="Hourly health check"
)The server is organized into modular components:
collaboration-tools/
├── src/
│ ├── main.py # MCP server entry point
│ ├── config.py # Configuration management
│ ├── browser_tools.py # Browser automation
│ ├── notification_tools.py # Email & IM notifications
│ ├── hitl_tools.py # Human-in-the-loop
│ └── timer_tools.py # Timer management
├── requirements.txt # Python dependencies
├── env.example # Example configuration
└── README.md # This file
- Python 3.12 for the root
ch4install (browser-userequires Python 3.11+) - OpenAI API key (for browser AI agent tasks)
- Optional: Email/IM service credentials
- Playwright browsers for browser automation
If browser automation fails:
# Reinstall Playwright browsers
playwright install chromium --force- For Gmail, use an App Password
- Ensure "Less secure app access" is NOT enabled (use App Passwords instead)
- Create a bot via @BotFather
- Get your chat ID from @userinfobot
If you see errors like "ChatOpenAI is not fully defined" or Pydantic validation errors:
- This is a known compatibility issue between LangChain and Pydantic v2
- The fix: ChatOpenAI is now initialized on-demand only when needed (in
browser_execute_task) - Simple browser navigation doesn't require OpenAI API key
- Only autonomous browser tasks (
browser_execute_task) requireOPENAI_API_KEY
MIT License
Contributions are welcome! Please feel free to submit issues or pull requests.