Skip to content
clusters: prooflayer · edgemarket · edgefinance · synthforge · mediakit · wordmint · webprobe · locale · comppoint · rollforge · bestiary · statline · matchpoint · retail · agentops · browserworkflow · modelrouter · compose
$ man video-summarize

/video-summarize

agentutility / mediakit / video-summarize
PRICE / CALL
$0.10
USDC · base mainnet · scheme: exact
METHOD
POST
CLUSTER
mediakit
CATEGORY
ai
STATUS
● live
NAME
video-summarize — summarizes videos, podcasts, and lectures in one call: whisper v3 transcribes, then mistral summarizes
SYNOPSIS
POST https://x402.agentutility.ai/video-summarize
     Content-Type: application/json
     X-PAYMENT:    <signed-transferWithAuthorization>

     { ... }
↳ first call → 402 Payment Required. Sign USDCtransferWithAuthorization, retry with theX-PAYMENT header.
DESCRIPTION

Summarizes videos, podcasts, and lectures in one call: Whisper v3 transcribes, then Mistral summarizes. 5 styles (tldr, bullets, paragraph, executive, chapters); returns summary + transcript; 60 min max. Use it as a video summarizer, podcast summarizer, or lecture notes generator.

INPUT — request schema
propertytypedescriptionreq?
media_urlstringURL of the video/audio file to transcribe and summarize.required
stylestringSummary format: tldr, bullets, paragraph, executive, or chapters. Optional, default bullets.
enum: tldr · bullets · paragraph · executive · chapters
optional
max_wordsnumberTarget word count for the summary. Optional, 1-1000.optional
languagestringSource language hint passed to Whisper (ISO code). Optional, auto-detected if omitted.optional
OUTPUT — response shape
fieldtypedescription
summarystringGenerated summary text in the requested style.
stylestringSummary style actually used (echoes the input style parameter).
transcriptstringFull Whisper v3 transcript of the source media.
transcript_charsnumberCharacter count of the returned transcript.
duration_secondsnumberLength of the source media in seconds.
detected_languagesarrayLanguages Whisper detected in the audio, as ISO codes.
summary_modelstringMistral model identifier used to write the summary.
transcribe_modelstringWhisper model identifier used for transcription (v3).
source_urlstringEcho of the media_url that was transcribed.
EXAMPLES — two ways to call
EXAMPLE 1 · curl
curl -X POST https://x402.agentutility.ai/video-summarize \
  -H 'Content-Type: application/json' \
  -d '{ }'
first response = 402 Payment Required with payment requirements; sign + retry with X-PAYMENT.
EXAMPLE 2 · mcp
# Install the MCP package for this endpoint's cluster
npx -y @agentutility/mcp-<cluster>

# Required: EVM private key with USDC on Base
export X402_PRIVATE_KEY=0x...

# Then call the video-summarize tool from your MCP-aware agent.
MCP server handles payment automatically — your coding agent just calls the tool by name.
METADATA
tags
videopodcastsummarizetranscribewhisper
env
FAL_KEY · VENICE_API_KEY
methods
POST
cluster
mediakit
price
$0.10 USDC per call
ADJACENT — other endpoints in mediakit
endpointdescriptionprice
doc-to-jsonConverts any document (PDF, DOCX, PPT, XLSX, or image) into structured JSON matching a caller-supplied schema.$0.10
extract-tablesDetects and extracts every table from a PDF document, returning structured JSON or CSV per table.$0.10
pdf-extract-tablesExtracts every table from a PDF, digital or scanned, and returns row-by-column text matrices page-by-page.$0.10
pdf-table-extractExtracts tables from digital or scanned PDFs, returning row/column matrices, CSV output, page numbers, and optional cell boxes.$0.10
pdf-table-extractorFinds tables in digital or scanned PDFs and returns row-by-column matrices, page numbers, and optional cell bounding boxes.$0.10
pdf-to-jpgConvert a PDF into page images for previews, thumbnails, or archiving.$0.10
speaker-diarizeTranscribe audio or video and label who said what.$0.10
transcribeConvert spoken audio or video into text.$0.10
SEE ALSO
agentutility · mediakit · x402 · mcp · llms.txt · registry.json · bazaar.x402.org