ClawTrust LogoClawTrust
PaddleOCR Document Parsing

PaddleOCR Document Parsing

by Bobholamovic · v1.0.0

Data Analysis
ClawHub
8.0
/ 10
1 evaluations
2.1k Downloads

Overview

Provides advanced OCR-powered document parsing via PaddleOCR, returning a rich, structured JSON representation of documents (text, tables, formulas, figures, layout, and reading order) by calling the `python scripts/vl_caller.py` wrapper script against a configured PaddleOCR API endpoint.

Key Advantages

1.Full document-structure extraction: text, tables, formulas (with LaTeX), charts/figures, headers/footers, page numbers, stamps, layout, and reading order.
2.Clear, strict workflow: always uses the provided `vl_caller.py` script with `--file-url` or `--file-path`, never does ad-hoc parsing or vision fallbacks, reducing inconsistency.
3.Well-defined JSON envelope with both high-level combined text and per-page structured outputs (`text`, `result[n].markdown`, `result[n].prunedResult`).
4.Strong guidance for configuration and troubleshooting, including handling missing API URL/token, authentication errors, quota limits, and unsupported formats.
5.Supports large and complex documents, including multi-column layouts and up to 100-page PDFs per request, with tooling for page subsetting via `split_pdf.py`.","Designed to always return complete, un‑

Use Cases

  • Parsing invoices, receipts, and financial reports that contain structured tables, totals, and mixed text/table layouts.
  • Extracting structured content from academic papers and technical reports, including mathematical formulas and references.
  • Processing multi-column documents like newspapers, magazines, and brochures where layout and reading order matter.
  • Digitizing complex PDFs with mixed content (text, tables, figures, seals) into a structured JSON format for downstream analysis or indexing.
  • Extracting all text, tables, and layout metadata from corporate documents, contracts, and reports for compliance or audit workflows.","Building downstream applications (search, analytics, RAG) that需要"

Evaluation Scores

8.0
/ 10
Reliability
7.0
Functionality
8.8
Usability
8.3
Safety
8.0
Performance
7.5
Compatibility
8.0

Based on 1 evaluation · Latest: 3/19/2026

Download Trend

Loading...

Evaluation History (1)

8.0/103/19/2026
▼
OS: linux-x64LLM: z-ai/glm-4.5-air
**Overall verdict:** Strong, specialized skill for rich document OCR and layout parsing, best suited for complex PDFs and images where tables, formulas, and multi-column layouts matter. It is tightly scoped, well-documented, and enforces correct use of the PaddleOCR API through a single script entrypoint. **Strengths** - High-fidelity structural extraction (text, tables, formulas, figures, layout, reading order) with a clean JSON envelope. - Clear operational contract: always use `python scripts/vl_caller.py`, never fall back to in-model vision, and always return complete content as requested. - Good guidance for configuration, API errors, and large-file handling, including page-range splitting. **Key risks / constraints** - Hard dependency on an external PaddleOCR API; failures (misconfiguration, quota, network) cannot be mitigated by the agent and must be resolved by the user. - Requires prior environment configuration for API URL, token, and optional timeout; initial setup may be non-trivial for non-technical users. - All documents are sent to a third-party OCR service, which may be a concern for sensitive or regulated content. **Recommended scenarios** - Complex document ingestion pipelines where you need full structure (tables, formulas, layout) rather than plain OCR text. - Back-office or research workflows processing invoices, financial statements, scientific papers, and multi-column PDFs. - As a dedicated OCR/layout microservice within a larger OpenClaw agent system, where another skill or agent performs downstream summarization, analysis, or search on the structured JSON output.

Comments (0)

Post a Comment

No comments yet. Be the first!