shotfun-creator
Health Warn
- License — License: NOASSERTION
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 8 GitHub stars
Code Warn
- process.env — Environment variable access in scripts/cli/doctor.js
Permissions Pass
- Permissions — No dangerous permissions requested
No AI report is available for this listing yet.
Agent skill collection for AI content production across images, video, audio, and digital humans
shotfun-creator
shotfun-creator 面向所有 AI 内容生产场景,是覆盖图片、视频、声音、数字人等能力的 skill 集合。它负责理解用户目标、自主选择合适的可用技能,并完成内容生产。用户只需要输入内容目标,它会帮助拆解任务、规划流程、完成工作任务,并可把已实现的工作流程沉淀为用户自己的 skill。
项目分为四层:
SKILL.md:主 skill,负责理解用户目标,并路由到合适的工作流、任务 skill 或原子服务。workflow-skills/:复杂工作流,例如默认口播内容生产。task-skills/:输入输出明确的任务能力,例如公众号封面、参考视频分析、口播视频、声音、数字人等内容生产任务。scripts/services/:稳定的 API 原子服务,涵盖生图、参考生视频、TTS/声音克隆、视频分析、ASR、字幕、素材管理和文本任务。
优先使用能完整覆盖目标的最高层能力;只有没有合适 workflow/task skill 时,才下钻到 CLI 或 atomic service。
安装和使用
推荐安装方式
这是一个标准的 agent skill 项目。最简单的方式是让你的 agent 客户端执行 skills CLI 安装命令,例如 Claude Code、Codex、Hermes、OpenClaw 或其他支持 skill 的客户端:
帮我安装这个 skill:npx skills add shotfun-ai/shotfun-creator
agent 会帮你通过 npx skills add shotfun-ai/shotfun-creator 安装到本地 skills 目录,并在需要时提示你重启或重新加载 skill。
手动安装
如果你想手动安装,请先选择当前客户端的 skills 目录。常见目录如下;如果你的客户端配置了其他目录,请以实际配置为准。
| 客户端 | 常见 skills 目录 |
|---|---|
| Codex | ~/.codex/skills |
| Claude Code | ~/.claude/skills |
| OpenClaw | ~/.openclaw/skills |
| Hermes | 使用 Hermes 配置的 skills 目录 |
然后执行:
export AGENT_SKILLS_DIR="<your-agent-skills-dir>"
mkdir -p "$AGENT_SKILLS_DIR"
git clone https://github.com/shotfun-ai/shotfun-creator.git "$AGENT_SKILLS_DIR/shotfun-creator" || \
(cd "$AGENT_SKILLS_DIR/shotfun-creator" && git pull --ff-only)
安装完成后,重启或重新加载你的 agent 客户端,让 shotfun-creator skill 生效。
配置 ShotFun API Key
大部分生成任务需要 ShotFun API Key。在安装目录创建 .env.local:
cd "$AGENT_SKILLS_DIR/shotfun-creator"
printf "SHOTFUN_API_KEY=your_api_key_here\nSHOTFUN_PROJECT_CODE=default\n" > .env.local
请不要公开 .env.local。该文件已被 git 忽略。
验证安装
重启或重新加载 agent 客户端后,可以直接问:
使用 shotfun-creator 列一下当前可用 skill
也可以做一次 CLI 检查:
cd "$AGENT_SKILLS_DIR/shotfun-creator/scripts"
npm run doctor
更新项目
后续更新:
cd "$AGENT_SKILLS_DIR/shotfun-creator"
git pull --ff-only
更新后请再次重启或重新加载 agent 客户端,因为 SKILL.md 和子 skill 文件通常在启动时加载。
当前可用 Skill
主 Skill
shotfun-creator:主入口。理解用户目标,并路由到合适的 workflow skill、task skill 或原子服务。
Workflow Skills
koubo:默认口播内容生产工作流。先用gpt-image2(即 gpt-image-2)生成/确认口播形象图,再准备脚本、声音参考、口播视频/片段包、可选封面和运行产物。
Task Skills
wechat-write-publish-allinone:生成公众号文章、封面图,并发布到公众号草稿箱。wechat-cover-image:生成微信公众号封面图。xhs-images-gen:生成小红书/RedNote 图片卡片。universal-content-to-image:把任意内容生成展示图片,例如产品促销图、培训说明图、信息图等。reference-video-analysis:分析参考视频,抽帧、识别节奏、总结视觉风格。talking-head-scene-image:根据要求、主播照片和可选场景图生成口播场景图。scripted-talking-video:统一的口播视频生成 skill,支持短脚本单镜头和多镜头口播/B-roll 视频包。hyperframes-project:创建、检查并可选渲染可编辑的 HyperFrames 视频项目。workbench-web-skill:为指定 skill run 和产物生成简单 Web 工作台。
原子能力
已同步 ShotFun Core 的通用原子服务、CLI、模型注册表和音色目录:
- 图片生成与编辑、参考生视频、语音合成、声音克隆和多模态音频生成。
- 视频分析(支持本地代理和分段)、音视频转录、SRT 生成、字幕烧录与字幕去除。
- 素材上传与报白、可用积分余额查询。
- 已有后端文本任务调用、批量执行与恢复、结构化产物保存、模型 JSON 修复审查及兼容视频拼接。
调用契约和参数见 references/model-catalog.md 与 references/calling-conventions.md。需要本地视频处理时安装 ffmpeg / ffprobe;会员区域默认 CN,国际账号设置 SHOTFUN_REGION=EN。参考声音和素材应由用户拥有或获得授权。
许可证
本项目采用 PolyForm Noncommercial License 1.0.0。
允许个人学习、研究、实验和非商业评估。商业使用需要获得 ShotFun 单独书面商业授权。
商业使用包括但不限于:用于商业产品、SaaS 或托管服务、客户交付、生产业务流程、付费内部服务、转售、再授权、白标或其他变现用途。
如需商业授权,请查看 COMMERCIAL_LICENSE.md。
shotfun-creator English Guide
shotfun-creator is a skill collection for AI content production across images, video, audio, digital humans, and related formats. It understands the user's content goal, autonomously selects the right available skills, completes the production task, and can turn implemented workflows into the user's own reusable skills.
The repository is organized into four layers:
SKILL.md: the main routing skill. It interprets the user's goal, chooses an available workflow skill, task skill, or atomic service, and reports final artifacts.workflow-skills/: curated multi-step workflows for complex outcomes, such as the default talking-head content workflow.task-skills/: reusable task-level skills with clear inputs and outputs, such as cover images, video download and analysis, talking-head videos, audio, digital humans, and other content production tasks.scripts/services/: stable atomic API services for images, reference videos, TTS and voice cloning, video analysis, ASR, subtitles, assets, and text tasks.
Use the highest layer that fully matches the user's intent. Drop down only when the higher layer does not exist or is too broad.
Install And Use
Quick Install
This repository is a standard agent skill package.
Recommended: ask your agent client, such as Claude Code, Codex, Hermes, OpenClaw, or another skill-capable agent, to run the skills CLI install command:
Install this skill: npx skills add shotfun-ai/shotfun-creator
The agent should install it with npx skills add shotfun-ai/shotfun-creator and restart/reload skills if needed.
Manual Install
If you prefer installing manually, choose the skills directory for your current agent client first. Common locations are listed below; if your client uses a custom path, use that configured path instead.
| Client | Common skills directory |
|---|---|
| Codex | ~/.codex/skills |
| Claude Code | ~/.claude/skills |
| OpenClaw | ~/.openclaw/skills |
| Hermes | the skills directory configured in Hermes |
Then run:
export AGENT_SKILLS_DIR="<your-agent-skills-dir>"
mkdir -p "$AGENT_SKILLS_DIR"
git clone https://github.com/shotfun-ai/shotfun-creator.git "$AGENT_SKILLS_DIR/shotfun-creator" || \
(cd "$AGENT_SKILLS_DIR/shotfun-creator" && git pull --ff-only)
Then restart or reload your agent client so the shotfun-creator skill is loaded.
Configure ShotFun API Key
Most generation tasks need a ShotFun API key. Create .env.local in the installed skill directory:
cd "$AGENT_SKILLS_DIR/shotfun-creator"
printf "SHOTFUN_API_KEY=your_api_key_here\nSHOTFUN_PROJECT_CODE=default\n" > .env.local
Keep .env.local private. It is ignored by git.
Verify
After restarting or reloading your agent client, ask:
List the currently available skills for shotfun-creator.
For a CLI smoke test:
cd "$AGENT_SKILLS_DIR/shotfun-creator/scripts"
npm run doctor
Update
To update later:
cd "$AGENT_SKILLS_DIR/shotfun-creator"
git pull --ff-only
Restart or reload your agent client again after updating, because SKILL.md and nested skill files are usually loaded at startup.
Available Skills
Main Skill
shotfun-creator: the main entry point. It understands the user's goal and routes to a workflow skill, task skill, or atomic service.
Workflow Skills
koubo: default talking-head content workflow. It first generates or confirms the presenter image withgpt-image2(also referred to as gpt-image-2), then prepares the script, voice reference, talking-head video or clip package, optional cover, and run artifacts.
Task Skills
wechat-write-publish-allinone: generate a WeChat article, cover image, and draft-box publishing package.wechat-cover-image: generate a WeChat article cover image.xhs-images-gen: generate Xiaohongshu/RedNote image cards.universal-content-to-image: turn arbitrary content into display images, product images, training explainer images, and similar visuals.reference-video-analysis: analyze reference videos, extract frames, prepare ASR/transcript artifacts, and summarize reusable visual style.talking-head-scene-image: generate a talking-head scene image.scripted-talking-video: unified talking video generation skill supporting short single-shot talking-head videos and longer multi-shot presenter/B-roll packages.hyperframes-project: create, inspect, and optionally render editable HyperFrames video projects.workbench-web-skill: build a simple web workbench for a specified skill run and its artifacts.
See CREDITS.md for external design-methodology acknowledgements used by specific workflows.
Atomic Capabilities
The reusable ShotFun Core services, CLIs, model registry, and voice catalog are synchronized into this project:
- Image generation and editing, reference-to-video generation, TTS, voice cloning, and multimodal audio generation.
- Video analysis with local proxies and chunking, audio/video transcription, SRT generation, subtitle burning, and subtitle removal.
- File uploads and asset registration, and available credit balance queries.
- Existing backend text tasks, resumable batches, contract-based artifact saving, model JSON repair review, and compatible video assembly.
See references/model-catalog.md and references/calling-conventions.md for contracts and parameters. Local video processing requires ffmpeg / ffprobe. The account region defaults to CN; set SHOTFUN_REGION=EN for international accounts. Reference voices and media must be owned or authorized by the user.
License
This project is source-available under the PolyForm Noncommercial License 1.0.0.
Personal learning, research, experimentation, and non-commercial evaluation are allowed. Commercial use requires a separate written commercial license from ShotFun.
Commercial use includes using this project or modified versions in commercial products, SaaS or hosted services, customer delivery, production business operations, paid internal services, resale, sublicensing, white-labeling, or other monetized scenarios.
For commercial permission, see COMMERCIAL_LICENSE.md.
Reviews (0)
Sign in to leave a review.
Leave a reviewNo results found