Skip to content
clusters: prooflayer · edgemarket · edgefinance · synthforge · mediakit · wordmint · webprobe · locale · comppoint · rollforge · bestiary · statline · matchpoint · retail · agentops · browserworkflow · modelrouter · compose
$ man pdf-extract-tables

/pdf-extract-tables

agentutility / mediakit / pdf-extract-tables
PRICE / CALL
$0.10
USDC · base mainnet · scheme: exact
METHOD
POST
CLUSTER
mediakit
CATEGORY
uncategorized
STATUS
live
NAME
pdf-extract-tables extracts every table from a pdf, digital or scanned, and returns row-by-column text matrices page-by-page
SYNOPSIS
POST https://x402.agentutility.ai/pdf-extract-tables
     Content-Type: application/json
     X-PAYMENT:    <signed-transferWithAuthorization>

     { ... }
↳ first call → 402 Payment Required. Sign USDCtransferWithAuthorization, retry with theX-PAYMENT header.
DESCRIPTION

Extracts every table from a PDF, digital or scanned, and returns row-by-column text matrices page-by-page. AI + OCR pipeline with optional cell bounding boxes for downstream layout reconstruction and an optional page_range filter ('1-5', '3', '1,3,5'). Handles merged headers, multi-page financial statements, balance sheets, lab results, scanned reports. 30 pages max. Sibling of pdf-to-markdown using the same Datalab backend, but pre-parsed to tables only. Use it as a PDF table extractor, scanned-table parser, financial-table OCR, multi-page table consolidator, or Datalab Marker tables endpoint.

INPUTrequest schema
propertytypedescriptionreq?
pdf_urlstringPublic URL of a PDF file (http or https). Must be directly fetchable, not behind auth or a viewer redirect. Max 30 pages.required
page_rangestringOptional 1-indexed page filter applied after extraction. Accepts ranges, single pages, or comma-lists: '1-5', '3', '1,3,5'. Default: all pages.optional
OUTPUTresponse shape
fieldtypedescription
source_urlstringEchoes back the PDF URL that was extracted, for traceability.
page_countstringTotal number of pages in the input PDF (capped at 30).
tablesstringArray of detected tables with page number, row × column text matrix, and optional cell bounding boxes.
sourcestringBackend identifier, typically 'datalab-marker', indicating the OCR/parsing engine used.
EXAMPLEStwo ways to call
EXAMPLE 1 · curl
curl -X POST https://x402.agentutility.ai/pdf-extract-tables \
  -H 'Content-Type: application/json' \
  -d '{ }'
first response = 402 Payment Required with payment requirements; sign + retry with X-PAYMENT.
EXAMPLE 2 · mcp
# Install the MCP package for this endpoint's cluster
npx -y @agentutility/mcp-<cluster>

# Required: EVM private key with USDC on Base
export X402_PRIVATE_KEY=0x...

# Then call the pdf-extract-tables tool from your MCP-aware agent.
MCP server handles payment automatically — your coding agent just calls the tool by name.
METADATA
tags
pdftable-extractionocrmediakitdocument-parsingfinancial-tablesdatalabpdf-tables
methods
POST
cluster
mediakit
price
$0.10 USDC per call
ADJACENTother endpoints in mediakit
endpointdescriptionprice
doc-to-jsonConverts any document (PDF, DOCX, PPT, XLSX, or image) into structured JSON matching a caller-supplied schema.$0.10
extract-tablesDetects and extracts every table from a PDF document, returning structured JSON or CSV per table.$0.10
pdf-table-extractExtracts tables from digital or scanned PDFs, returning row/column matrices, CSV output, page numbers, and optional cell boxes.$0.10
pdf-table-extractorFinds tables in digital or scanned PDFs and returns row-by-column matrices, page numbers, and optional cell bounding boxes.$0.10
pdf-to-jpgConverts a PDF to JPG, PNG, or WEBP images, rendering every page at configurable DPI (36-600) and returning one image URL per page.$0.10
speaker-diarizeSpeaker diarization / who-said-what transcription.$0.10
transcribeTranscribe video to text.$0.10
video-summarizeSummarizes videos, podcasts, and lectures in one call: Whisper v3 transcribes, then Mistral summarizes.$0.10
SEE ALSO
agentutility · mediakit · x402 · mcp · llms.txt · registry.json · bazaar.x402.org