8.0
/ 10
1 evaluations
2.1k Downloads
Overview
Provides advanced OCR-powered document parsing via PaddleOCR, returning a rich, structured JSON representation of documents (text, tables, formulas, figures, layout, and reading order) by calling the `python scripts/vl_caller.py` wrapper script against a configured PaddleOCR API endpoint.
Key Advantages
1.Full document-structure extraction: text, tables, formulas (with LaTeX), charts/figures, headers/footers, page numbers, stamps, layout, and reading order.
2.Clear, strict workflow: always uses the provided `vl_caller.py` script with `--file-url` or `--file-path`, never does ad-hoc parsing or vision fallbacks, reducing inconsistency.
3.Well-defined JSON envelope with both high-level combined text and per-page structured outputs (`text`, `result[n].markdown`, `result[n].prunedResult`).
4.Strong guidance for configuration and troubleshooting, including handling missing API URL/token, authentication errors, quota limits, and unsupported formats.
5.Supports large and complex documents, including multi-column layouts and up to 100-page PDFs per request, with tooling for page subsetting via `split_pdf.py`.","Designed to always return complete, un‑
Use Cases
- Parsing invoices, receipts, and financial reports that contain structured tables, totals, and mixed text/table layouts.
- Extracting structured content from academic papers and technical reports, including mathematical formulas and references.
- Processing multi-column documents like newspapers, magazines, and brochures where layout and reading order matter.
- Digitizing complex PDFs with mixed content (text, tables, figures, seals) into a structured JSON format for downstream analysis or indexing.
- Extracting all text, tables, and layout metadata from corporate documents, contracts, and reports for compliance or audit workflows.","Building downstream applications (search, analytics, RAG) that需要"
Evaluation Scores
8.0
/ 10
Reliability
7.0
Functionality
8.8
Usability
8.3
Safety
8.0
Performance
7.5
Compatibility
8.0
Based on 1 evaluation · Latest: 3/19/2026
Download Trend
Loading...
Evaluation History (1)
8.0/103/19/2026▼
OS: linux-x64LLM: z-ai/glm-4.5-air
**Overall verdict:** Strong, specialized skill for rich document OCR and layout parsing, best suited for complex PDFs and images where tables, formulas, and multi-column layouts matter. It is tightly scoped, well-documented, and enforces correct use of the PaddleOCR API through a single script entrypoint.
**Strengths**
- High-fidelity structural extraction (text, tables, formulas, figures, layout, reading order) with a clean JSON envelope.
- Clear operational contract: always use `python scripts/vl_caller.py`, never fall back to in-model vision, and always return complete content as requested.
- Good guidance for configuration, API errors, and large-file handling, including page-range splitting.
**Key risks / constraints**
- Hard dependency on an external PaddleOCR API; failures (misconfiguration, quota, network) cannot be mitigated by the agent and must be resolved by the user.
- Requires prior environment configuration for API URL, token, and optional timeout; initial setup may be non-trivial for non-technical users.
- All documents are sent to a third-party OCR service, which may be a concern for sensitive or regulated content.
**Recommended scenarios**
- Complex document ingestion pipelines where you need full structure (tables, formulas, layout) rather than plain OCR text.
- Back-office or research workflows processing invoices, financial statements, scientific papers, and multi-column PDFs.
- As a dedicated OCR/layout microservice within a larger OpenClaw agent system, where another skill or agent performs downstream summarization, analysis, or search on the structured JSON output.
Comments (0)
No comments yet. Be the first!