ocr-service
High-precision Optical Character Recognition (OCR) service. Supports text detection and extraction from multi-language, multi-format images, and provides text area coordinates and confidence scores, s…
Browse curated skills with source links, package snapshots, README assets and install signals in one calm, searchable catalog.
High-precision Optical Character Recognition (OCR) service. Supports text detection and extraction from multi-language, multi-format images, and provides text area coordinates and confidence scores, s…
用 HTML/CSS 生成内容配图(公众号封面 + 概念对比卡 + 文内图),输出 PNG。当用户说「给这篇配封面」「生成封面」「公众号头图」「给这段做个对比图」「概念对照卡」「文内配图」时使用。
帮助内容创作者、IP 与角色设计团队、种草博主和视觉策划直接完成“GPT Image 2 角色一致性图片”:上传一张现有图片并说明要保留和改变的内容,就能完成重绘、精修或风格调整;通过 AI Hive 使用时,生成前自动上传参考图,提交后自动保存任务、查询进度并下载图片。适用于电商主图、商品详情页、广告 KV、海报、带货、种草、社媒配图、产品精修、换背景与角色一致性内容。 Use this ski…
RNN+Transformer hybrid with O(n) inference. Linear time, infinite context, no KV cache. Train like GPT (parallel), infer like RNN (sequential). Linux Foundation AI project. Production at Windows, Offi…
Converts PPT/slides with narration and subtitles into a video. It extracts/renders slide images from PPTX files or slides.json, generates narration using TTS, and adds subtitles to create the final MP…
Research alchemist, biomedical expert, and research narrative. Analyze research charts, generate in-depth narratives combined with literature from PubMed, etc. Requires coordination with MCP tools; us…
Automatic exploratory data analysis and visualization with a single line of code - generates comprehensive charts, detects patterns, and exports to HTML/notebooks
帮助品牌市场、广告创意、平面设计和社交媒体运营团队直接完成“Nano Banana Pro 精准文字图片”:既可从文字生成,也可加入参考图帮助控制主体、构图和视觉风格;通过 AI Hive 使用时,生成前自动上传参考图,提交后自动保存任务、查询进度并下载图片。适用于电商主图、商品详情页、广告 KV、海报、带货、种草、社媒配图、产品精修、换背景与角色一致性内容。 Use this skill for…
Research latest ComfyUI models, techniques, and community discoveries. Monitors YouTube channels, GitHub repos, and HuggingFace. Updates reference files with timestamped findings and flags stale infor…
生成、合成并编辑产品图、品牌视觉、海报、社交媒体配图、插画和照片的不同版本。
帮助电商运营、商品摄影、品牌商品团队和直播带货团队直接完成“Nano Banana 2 商品详情页”:既可从文字生成,也可加入参考图帮助控制主体、构图和视觉风格;通过 AI Hive 使用时,生成前自动上传参考图,提交后自动保存任务、查询进度并下载图片。适用于电商主图、商品详情页、广告 KV、海报、带货、种草、社媒配图、产品精修、换背景与角色一致性内容。 Use this skill for Na…
Generate images using the Doubao SeeDream API based on text prompts. Use this skill when users request AI-generated images, artwork, illustrations, or visual content creation. The skill handles API ca…
Look up 450+ Japanese cinematic, photography, and production terms with English equivalents and prompt-ready phrases for Seedance 2.0 across 20 categories, including filter-safe vocabulary for action,…
To be used when the user explicitly requests to 'optimize the prompt', 'improve the prompt', 'polish the instructions', or 'convert a basic prompt into a best-practice version'. Based on official best…
Data analyst. Supports pandas, numpy, visualization, and statistical analysis.
本地语音转写保险柜与转写协调 Skill。把会议录音、语音笔记、采访、客服通话、语音备忘、讲座等本地音频,用 SenseVoice/FunASR 经 OpenVINO 在本机(CPU/GPU/NPU)转成文字;原始音频和未脱敏转写不进远程大模型;转写整理成带时间戳、置信度、敏感串候选的结构化转写包,下游纪要/字幕/填表交给现成文档/字幕/表格工具,最后给出已转写/低置信度/需脱敏/需确认的验收报告…
ASR (Automatic Speech Recognition) — enhanced speech-to-text built on Doubao large model, with audio preprocessing, denoising, and extended analysis capabilities. Async API. Choose this skill when: - …
全能音乐创作助手——交互式选曲风格(说唱、R&B、儿歌、流行、摇滚、电 子、爵士、古典、民谣、鬼畜音MAD等40+子风格),从歌词创作到HappyShrimp风格提示词 一站式输出。支持中文/英文/双语/日语。当用户要求写歌曲、创作音乐、做儿歌、写Rap、R&B、流行歌曲、编曲、作曲、鬼畜、音MAD时使用。
INVOKE THIS SKILL when optimizing, improving, or debugging LLM prompts using production trace data, evaluations, and annotations. Covers extracting prompts from spans, gathering performance signal, an…
Implement ReasoningBank adaptive learning with AgentDB's 150x faster vector database. Includes trajectory tracking, verdict judgment, memory distillation, and pattern recognition. Use when building se…
Search language, e.g. zh-CN, en-US
Use when selecting AI models, configuring API parameters, or implementing LLM calls. Covers OpenAI (GPT-5.2, GPT-5.1, GPT-4.1, o3), Anthropic (Claude 4.5), Google (Gemini 2.5/3), DeepSeek (V3.2, R1), …
Read and comprehend the novel text, summarizing it into a smooth story outline. Suitable for initial novel screening and generating a 500-800 word story outline.
Process large document corpora (1000+ docs, millions of tokens) through knowledge graph construction and stateful multi-hop reasoning. Use when (1) User provides a large corpus exceeding context limit…