7.5
/ 10
1 evaluations
1.9k Downloads
Overview
Multi-model AI image and video generation via the Yollomi API, exposed through a single unified OpenClaw tool for both text-to-image/video and image-editing workflows.
Key Advantages
1.Unified endpoint and tool for many image and video models (Flux, SD, VEO, Sora, Runway, etc.), simplifying integration.
2.Covers both creation and editing: text-to-image, text-to-video, background removal, object remover, face swap, upscaling, restoration, virtual try-on, and more.
3.Dedicated listModels tool requiring no auth, allowing dynamic discovery of supported models and their credit costs at runtime.
4.Simple auth and configuration via environment variables (YOLLOMI_API_KEY and optional YOLLOMI_BASE_URL).
5.Explicit support for aspect ratios and model-specific parameters, enabling fine-grained control over outputs and costs.
Use Cases
- Chatbots or agents that need to generate illustrative images for user prompts (e.g., product mockups, concept art, social media graphics).
- Video-generating assistants that produce short clips or explainer content from text prompts using third-party video models.
- Creative tools for background removal, object removal, and face swapping in user-uploaded photos.
- E-commerce or fashion assistants offering virtual try-on and background generation for product imagery.
- Photo enhancement workflows such as upscaling, photo restoration, and image editing with text instructions (e.g., Qwen image edit).
Evaluation Scores
7.5
/ 10
Reliability
7.0
Functionality
9.0
Usability
8.0
Safety
5.5
Performance
7.5
Compatibility
8.5
Based on 1 evaluation · Latest: 3/19/2026
Download Trend
Loading...
Evaluation History (1)
7.5/103/19/2026▼
OS: darwin-arm64LLM: google/gemini-2.5-flash
**Quick judgement:** A capable, feature-rich multi-model media generation skill that cleanly exposes the Yollomi image/video API through two tools (`yollomi.generate`, `yollomi.listModels`). Strong choice when you explicitly want to tap into a broad set of third-party image and video models via one interface, but it inherits all the usual external-API and safety limitations.
**What it’s good for**
- Centralizing access to many SOTA image/video models (Flux, Imagen, SD, Veo, Sora, Runway, etc.) behind one tool.
- Building assistants that can both create and edit images (background removal, object remover, face swap, upscaling, restoration, virtual try-on).
- Cost-aware agents that can query `listModels` to pick models based on credit price and capabilities.
- Workflows needing configurable aspect ratios and multiple outputs from a single prompt.
**Key risks and limitations**
- **External dependency & credits:** All functionality depends on Yollomi’s availability, rate limits, and credit balance; 401/402 errors must be handled by the caller.
- **Latency and throughput:** Actual performance is determined by Yollomi and the underlying model providers; may be slow for heavier models (e.g., high-end video or "pro" image models).
- **Safety & content controls:** No explicit mention of built-in moderation or safety filters. If you need strict content controls (NSFW, violence, deepfakes, etc.), you must add your own policy checks and guardrails around this skill.
- **Model-specific parameters:** Different models require different arguments (e.g., `imageUrl`, `mask`, `clothImage`, `personImage`, `inputs` for video). Incorrect parameterization can lead to failures unless the calling agent is careful or uses `listModels`/docs.
**Recommended scenarios**
Use this skill when you:
- Want a single, unified way to call many image and video generation/editing models.
- Are comfortable relying on an external paid API (with keys and credits managed out-of-band).
- Can layer your own safety, validation, and error-handling logic around the tool.
Avoid it (or wrap it heavily) in contexts where you need guaranteed offline operation, strict content moderation, or fully predictable latency.
Comments (0)
No comments yet. Be the first!