# API Reference — MCP Scraper

Source: https://mcpscraper.dev/docs

API & SDK reference

# Build with every MCP Scraper tool.

One generated contract exposes all 364 MCP tools through four SDK packages, the CLI, direct cURL, and the hosted REST API.

View the SDK on GitHub ↗ Browse the tool surface Connection quickstart

Quick start

```
npm install mcpscraper-sdk
```

364 tools 4 SDK packages 30 typed namespaces

X-Ray attribution

## Measure the journey. Prove the outcome.

Install one tag without giving X-Ray repository access, compare seven observed attribution models, preserve campaign evidence, and keep customer-reported influence in a separate inspectable ledger.

### Product overview

Understand what the tag can control, why browser intent is not revenue, and which capabilities are currently available.

Explore X-Ray →

### Attribution methodology

Review all seven models, 40/20/40, time decay, evidence tiers, explicit reported weights, and honest view limits.

Read the methodology →

### SaaS measurement foundation

Follow six verifiable stages from tag coverage through consent and a persisted observed-event test.

Open the SaaS guide →

Before the first connection

## Know the monthly fee and the usage meter.

A paid plan supplies the shared Credit balance. Each active Nango-backed provider account is a separate 15,000-credit/month add-on drawn from that same balance, and connected work draws from it too as it runs.

1. Choose a paid plan
   Starter, Growth, and Scale include Credits, but no active provider connections.
2. Approve the connection fee
   Each active Nango connection charges 15,000 Credits from your balance per month. If you later disconnect it, a same-integration replacement inherits that paid period through the end of the billing month.
3. Use the shared Credits
   Connected work costs 2 Credits per function run, 2 per proxy request, and 5 per compute-second measured from milliseconds.
4. Review and remove
   Use the dashboard Credits tab for current rates, History for settled usage, and Integrations to disconnect an account. The completion message confirms its replacement billing rule.

For a scheduled connection job, add the 75-Credit base for each started run. Agent mode also adds 1.5× the model provider's reported cost. [See the billing example and full rate card.](/pricing#connected-billing)

## Two result pages with first-page PAA

Set `pages: 2` on `harvest_paa`, `harvest_paa_start`, or a REST harvest request to capture page 2 organic results before expanding People Also Ask on the original first page. Omit it for the one-page default. Page 2 does not add a second PAA graph.

```
{
  "query": "commercial truck insurance",
  "pages": 2,
  "maxQuestions": 10
}
```

Google may not offer a second page. A page-two timeout or unavailable page preserves available first-page results and questions. MCP responses expose `pagination`; raw REST results expose `diagnostics.pagination`. Check `requestedPages`, `capturedPages`, and `page2Status` before treating both pages as captured. Older saved results may omit this metadata or return `null`. Existing harvest pricing is unchanged.

## Authentication

Pass your key in the `x-api-key` header. Query-string API keys are rejected so keys do not end up in URLs, browser history, or logs.

```
x-api-key: sk_live_••••
```

Gmail · selection first

## Read the whole email, preserve its files, then act on the exact set you reviewed

Gmail uses named tools instead of making an agent assemble a bulk workflow from provider pages. Search and complete reads are separate from mailbox mutations; attachment references resolve to actual private files, not metadata-only placeholders.

### Complete reads

`gmail_search_messages` returns bounded previews. `gmail_get_message` returns the decoded body, headers, MIME parts, completeness flags, and opaque attachment references. Oversized content moves to a complete private artifact instead of being silently truncated.

### Files you can open

`gmail_get_attachment` fetches the original bytes behind an opaque reference and returns a private artifact plus a safe text window when supported. Direct Gmail read artifacts are retained for seven days.

### One reviewed set

`gmail_prepare_selection` freezes a query or 1–5,000 explicit message IDs for 24 hours. `gmail_export_selection`, `gmail_bulk_manage_messages`, `gmail_bulk_delete_messages`, and Memory planning consume that immutable owner-bound selection.

### 1. Search and freeze

```
mcpscraper tools call gmail_search_messages \
  --args '{"connectionId":"conn_…","query":"has:attachment after:2026/01/01","limit":25}' --json

mcpscraper tools call gmail_prepare_selection \
  --args '{"connectionId":"conn_…","query":"has:attachment after:2026/01/01"}' --json
```

### 2. Review the Memory plan

```
mcpscraper tools call gmail_prepare_memory_import \
  --args '{"selectionId":"gmail_selection_…","filingPolicy":"source_archive","destination":{"mode":"auto"},"attachmentPolicy":"preserve_all"}' --json
```

### 3. Start or resume

```
mcpscraper tools call gmail_import_to_memory \
  --args '{"importPlanId":"gmail_import_plan_…","idempotencyKey":"mailbox-archive-2026-q1"}' --json

mcpscraper tools call gmail_import_status \
  --args '{"ingestId":"gmail_ingest_…"}' --json
```

### Reversible mailbox changes

```
mcpscraper tools call gmail_bulk_manage_messages \
  --args '{"selectionId":"gmail_selection_…","expectedCount":42,"operation":"archive","idempotencyKey":"archive-reviewed-42"}' --json
```

### Library is the source default

`source_archive` writes one deterministic Library note per message, preserves reviewed attachments, and finishes with a Library manifest. `relationship_communications` is opt-in and files only exact existing People or Organization matches.

### Stored is not always indexed

`preserve_all` is the default attachment policy. Supported text, HTML, JSON, CSV, PDF, and DOCX can be extracted and indexed; other safe files can report `stored_not_indexed`. Refused files remain visible in the plan or manifest.

### Partial work is resumable

Status is read-only. A `partial` or interrupted import keeps completed receipts and reports whether to wait, retry, resume, or review. Calling `gmail_import_to_memory` again with the same plan and idempotency key resumes without duplicating records.

A Memory plan accepts at most 5,000 messages, 10,000 attachments, 25 MiB per attachment, 250 MiB of selected attachment bytes, and 2,000 Memory source parts. The planner refuses an unbounded import before the first write. It never creates inferred People, Organizations, Tasks, Deals, Projects, or deadlines.

Turning on account actions makes reviewed action tools eligible; it does not authorize a particular mutation. Trash and permanent deletion are different tools. Permanent deletion requires the unchanged selection, expected count, a literal confirmation, and the client's action confirmation flow.

Connected account → RAG

## Store one exact approved read as a searchable snapshot

`import_service_connection_to_memory` composes a tenant-owned connection read with a Memory write. The server rechecks the exact live `readTools` grant, redacts the result, marks it as untrusted provider data, writes it at a stable server-generated path, and indexes it in the same request.

### Import one Google Drive read

```
mcpscraper tools call import_service_connection_to_memory \
  --args '{"connectionId":"conn_…","providerConfigKey":"google-drive","tool":"<exact readTools entry>","args":{},"vault":"Library","title":"Drive snapshot"}' \
  --json
```

### Discover before calling

```
mcpscraper tools call list_service_connections --args '{}' --json
mcpscraper tools call describe_service_connection_tool \
  --args '{"connectionId":"conn_…","tool":"<readTools entry>"}' \
  --json
```

### Required authority

Supply the exact `connectionId`, matching `providerConfigKey`, one current read tool, and an existing ordinary notes-vault handle. `args` and `title` are optional.

### Stable and inspectable

The caller cannot choose the path. The same connection, tool, and canonical arguments upsert the same Markdown snapshot. The receipt includes path, hash, source bytes, indexed chunks, and search status.

### Bounded by design

Arguments are capped at 64 KiB and the serialized result at 1 MB. Binary, base64, oversized, inactive, unapproved, action, admin, channel-vault, and secure-vault requests fail closed.

This is one snapshot from one approved read. It does not paginate, hydrate a corpus, import an entire account, run continuously, propagate deletions, or build normalized tables. A successful response reports `search_ready`; if storage succeeds without indexed chunks, it explicitly reports `stored_not_indexed`. For a reviewed Gmail corpus with complete messages and attachments, use the named selection and import workflow above. `export_connected_service_data` remains the compatibility export for supported Gmail, Google Calendar, Google Search Console, Zoom, Resend, and Meta datasets.

Connected accounts

## Export a supported time range in one call

Use `export_connected_service_data` for its dedicated Gmail, Google Calendar, Google Search Console, Zoom, Resend, or Meta adapters. It is separate from the one-read Memory snapshot tool. For these supported datasets, the server owns pagination and delivery; large results become private JSONL instead of overflowing the agent context.

### Last seven days of email

```
mcpscraper tools call export_connected_service_data \
  --args '{"connectionId":"conn_…","dataset":"emails","lastDays":7}' \
  --json
```

### Renew an expired download link

```
mcpscraper tools call renew_connected_data_download \
  --args '{"artifactId":"artifact_…"}' \
  --json
```

### Resend operating history

```
mcpscraper tools call export_connected_service_data \
  --args '{"connectionId":"conn_…","dataset":"resend_data","lastDays":7,"maxItems":2000}' \
  --json
```

### Search Console performance

```
mcpscraper tools call export_connected_service_data \
  --args '{"connectionId":"conn_…","dataset":"search_console_performance","lastDays":28,"maxItems":5000}' \
  --json
```

### Inspect multiple Search Console URLs

```
mcpscraper tools call read_service_connection \
  --args '{"connectionId":"conn_…","tool":"inspect-urls","args":{"siteUrl":"sc-domain:example.com","urls":["https://example.com/","https://example.com/pricing"]}}' \
  --json
```

### Filtered scheduled Search Console data

```
mcpscraper tools call export_search_console_table_data \
  --args '{"tableName":"gsc_performance_…","filters":[{"column":"date","op":"gte","value":"2026-07-01"},{"column":"query","op":"like","value":"roof repair"}],"sort":[{"column":"clicks","direction":"desc"}],"maxRows":10000}' \
  --json
```

Small live exports return inline. Large exports are private, retained for seven days, and use 15-minute signed URLs. Each live invocation returns at most 5,000 records. Search Console's `search_console_performance` dataset walks every accessible property for a fresh requested range. For recurring history, create a scheduled action in `connection_sync` mode; after its first successful run, read the connection's `tableName`, use `table-describe` and `table-query` for exact filtering, then use `export_search_console_table_data` to download up to 50,000 matching stored rows. Search Console returns provider-selected top rows and is not guaranteed exhaustive. Resend's aggregate `resend_data` mode walks 12 collections; six core resources can also be selected individually, and each provider page hydrates at most 25 records. A partial live export returns a continuation object that preserves the original range.

Search Console also exposes API-only batches without creating a table: `inspect-urls` and `query-search-analytics-batch` are reads; `add-sites-batch`, `submit-sitemaps-batch`, `delete-sites-batch`, and `delete-sitemaps-batch` are gated actions. They return per-item receipts and can be granted to scheduled agent runs. Delete batches plan by default and require the exact confirmation token from `describe_service_connection_tool` before execution.

Official remote MCP · OAuth

## Connect Resend without adding 85 global tools

The dashboard opens Resend's hosted OAuth flow at `mcp.resend.com/mcp`. After consent, MCP Scraper discovers the official account-scoped tools and exposes only the reviewed intersection through its generic connection bridge and each schedule's exact grant.

85 official tools discovered

33 approved reads

45 gated actions

7 admin tools blocked

### Available to reads

Sent and received email, attachments, logs, contacts, broadcasts, templates, automations, domains, segments, topics, webhooks, and other account metadata.

### Available only behind action gates

Drafting, sending, scheduling, updating, verification, removal, event triggers, and audience changes require both the account action switch and the exact tool grant.

### Never delegated

API-key and OAuth-grant administration, the raw editor connector, and raw webhook creation remain outside MCP and Mastra execution. Credential-bearing fields are redacted from approved outputs.

Resend tools are connection-scoped capabilities, not 85 new globally callable names in the base MCP catalog. Inspect the current connection first, then grant an exact read or action. Policy source: [official Resend MCP repository ↗](https://github.com/resend/resend-mcp). OAuth and API-key fallback behavior: [official MCP documentation ↗](https://resend.com/docs/mcp-server).

Unified tool contract

## All 364 tools, on every supported surface

The catalog contains 238 scraper, connection, and automation tools plus 126 memory tools. Every surface is generated from the same checked-in manifest, so a missing or renamed tool fails the SDK parity check before release.

364 MCP tools

4 SDK packages

30 typed namespaces

| Surface | Install | 364-tool access | Reference |
|---|---|---|---|
| **Node.js — scraper** | `npm install mcpscraper-sdk` | `ScraperClient.tools` | [npm package ↗](https://www.npmjs.com/package/mcpscraper-sdk) |
| **Node.js — memory** | `npm install mcpscraper-memory-sdk` | `McpToolsClient` | [npm package ↗](https://www.npmjs.com/package/mcpscraper-memory-sdk) |
| **Python — scraper** | `pip install "mcpscraper-sdk @ git+https://github.com/VilovietaSEO/mcpscraper-sdk.git#subdirectory=packages/scraper-python"` | `ScraperClient.tools` | [source package ↗](https://github.com/VilovietaSEO/mcpscraper-sdk/tree/main/packages/scraper-python) |
| **Python — memory** | `pip install "mcpscraper-memory-sdk @ git+https://github.com/VilovietaSEO/mcpscraper-sdk.git#subdirectory=packages/memory-python"` | `McpToolsClient` | [source package ↗](https://github.com/VilovietaSEO/mcpscraper-sdk/tree/main/packages/memory-python) |
| **CLI** | `npm install -g mcpscraper-cli` | `tools list · describe · call` | [npm package ↗](https://www.npmjs.com/package/mcpscraper-cli) |
| **cURL** | `curl + jq` | `JSON-RPC tools/list + tools/call` | [364-tool catalog ↗](https://github.com/VilovietaSEO/mcpscraper-sdk/blob/main/docs/curl-tools.md) |

### CLI discovery and invocation

```
export MCPSCRAPER_API_KEY=sk_live_your_key
mcpscraper tools list
mcpscraper tools describe import_service_connection_to_memory
mcpscraper tools call credits_info --args '{}' --json
```

### Direct JSON-RPC with cURL

```
curl --retry 3 --retry-all-errors --retry-delay 1 \
  -X POST https://mcpscraper.dev/mcp \
  -H "x-api-key: $MCPSCRAPER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}'
```

Canonical source: [mcp.tools.json ↗](https://github.com/VilovietaSEO/mcpscraper-sdk/blob/main/contracts/mcp.tools.json). The SDK repository also contains generated types, package-specific examples, parity validators, and the complete cURL catalog.

## Endpoint map

Use the smallest surface that matches the job. SERP harvest is for Google results; website endpoints fetch public URLs; YouTube endpoints go directly to YouTube search, channels, and transcripts.

For the MCP `search_serp` tool, set `mode: "full"` on an unfiltered one-page search to request available same-page local-pack data. Setting `includeLocalPack` alone does not enable it. A populated CID list does not guarantee local-pack rows. Maps search returns candidates, and `websiteUrl: null` does not prove a business has no website. Hydrate selected candidates through `/maps/place` when complete profile fields are needed; each place lookup has its own charge.

| Method | Path | Use | Limit |
|---|---|---|---|
| POST | /mcp | Unified MCP JSON-RPC transport for all 364 scraper, browser, workflow, connection, billing, and memory tools | x-api-key; tools/list + tools/call |
| POST | /harvest/sync | Google SERP harvest: PAA, AI Overview, local pack, forums, videos, perspectives, entity IDs | up to 400s, max 100 questions |
| POST | /harvest | Async SERP harvest with job polling and optional webhook callback | Callback must be HTTPS public URL |
| GET | /jobs, /jobs/:id | List or fetch saved SERP harvest jobs for the authenticated user | User-scoped |
| POST | /extract-url | Synchronous compatibility route for one public URL extraction | Public http/https hosts only |
| POST | /extract-url/start | Start a durable single-URL extraction and return its job receipt | Use the same Idempotency-Key only for the same intended extraction |
| GET | /extract-url/status/:id | Read owner-scoped durable extraction status and terminal result | Use the returned job ID |
| POST | /map-urls | BFS internal URL discovery for a public website | maxUrls up to 10,000 |
| POST | /extract-site | Start a durable crawl that retains complete per-page JSON, acquired HTML, and Markdown | maxPages up to 10,000; retry with the same idempotency key only for the same intended crawl |
| GET | /extract-site/status/:id | Poll discovered, attempted, successful, failed, and remaining counts plus artifacts or a stable public error envelope | User-scoped |
| POST | /extract-site/read | Read a site-export manifest or bounded byte windows of one page's JSON, acquired HTML, or Markdown | User-scoped; max 1 MB per call; continue from nextOffset |
| POST | /extract-site/image | Read one preserved image by the imageId returned in the site-export manifest | User-scoped; images up to 10 MB |
| POST | /screenshot | Full-page screenshot of a public URL (desktop or mobile viewport) | Public http/https hosts only |
| POST | /serp-intelligence/capture | Structured SERP Intelligence snapshot for a query (product capture path) | 4 Credits headless or 14 headful; optional snapshots add 1 per attempted URL |
| POST | /serp-intelligence/page-snapshots | Capture optional public ranking-page evidence for one or more URLs/targets | 1 Credit per attempted URL; requested capacity held, unused capacity refunded; no private/localhost/internal URLs |
| POST | /youtube/harvest | YouTube search or channel harvest with video metadata and thumbnails | maxVideos 1-500 API, dashboard slider 10-200 |
| POST | /youtube/transcribe | Transcript by video ID with text, timestamp chunks, Markdown, and HTML | Depends on caption/Whisper availability |
| POST | /facebook/reels-inventory | Collect public profile or Page Reel URLs, then transcribe selected URLs | Up to 60 anonymous URLs; partial results and resumeUrls supported; 4 credits per scan, refunded with no new URLs |
| POST | /facebook/search | Discover advertisers by name or keyword in the Facebook Ad Library | maxResults 1-20 |
| POST | /facebook/page-intel | Scan all ads for an advertiser — copy, headline, CTA, status, creative type | Accepts libraryId, pageId, or query |
| POST | /facebook/ad | Extract full metadata for a single ad by library ID | Returns video, image, and text fields |
| POST | /facebook/video-transcribe | Transcribe an organic Facebook reel, video, watch, or share URL | quality best, hd, or sd; renders page first |
| POST | /tiktok/video-transcribe | Transcribe one public TikTok video or short share URL with timed chunks | 200-credit base plus 2 per minute for speech transcription; caption-only fallback is free |
| POST | /facebook/transcribe | Transcribe audio from a direct Facebook ad video CDN URL returned by /facebook/page-intel | Use the videoUrl from page-intel, not a public post URL |
| POST | /facebook/media | Download video or image from fbcdn.net CDN URLs | fbcdn.net and cdninstagram.com only |
| POST | /maps/search | Costs 5 Credits. Returns up to 50 candidates. websiteUrl can be null and ordinary search does not open every profile; use /maps/place selectively for complete profile fields. | maxResults 1-50; selectively hydrate candidates through /maps/place |
| POST | /maps/place | Start a recoverable Google Maps business profile lookup with declared fields | Idempotency-Key recommended; include: ['all'] or individual fields; reviews billed per saved card on completion |
| GET | /maps/place/runs/:runId | Read saved profile fields and run, field, and billing status | Owner scoped; first 50 reviews and 100 image references |
| GET | /maps/place/runs/:runId/reviews | Page through saved review cards | Opaque cursor; up to 50 per page |
| GET | /maps/place/runs/:runId/images | Page through saved image references | Opaque cursor; up to 100 per page |
| POST | /maps/place/runs/:runId/resume | Resume a partial or reconciled interrupted run | New Idempotency-Key required; previous billing must be reconciled |
| POST | /directory/run | Defaults to a durable background job. Poll its statusUrl without starting another search. background:false runs synchronously. Each failed city is refunded. | Idempotency-Key recommended; background:false runs synchronously |
| GET | /directory/jobs/:jobId | Returns progress, terminal city results, billing settlement, and any CSV artifact without starting or billing another search. | Owner scoped; polling does not bill another search |
| POST | /instagram/profile-content | Discover public Instagram post, reel, and tv links from a profile | maxItems + maxScrolls; public content only |
| POST | /instagram/media-download | Download text, images, reel audio/video, and optional transcript for one post | public / login-gated limits apply |
| POST | /reddit/thread | Capture a Reddit post and its full comment tree (author, score, depth, body) from any reddit.com thread URL | maxComments optional; handles Reddit bot wall via residential proxy |
| POST | /video/analyze | Start an async frame-by-frame + transcript video breakdown; $1 per 120 frames requested (max 480), videos up to 30 minutes | Direct .mp4/.webm/.mov URLs; returns runId |
| POST | /video/status | Poll a video breakdown run; returns progress, then the full report; reconciles billing down if fewer frames were usable | Free to call |
| GET | /workflows/definitions | List runnable workflows for agent packets, market audits, Maps/SERP comparison, PAA briefs, and AI Overview language guidance | User-scoped |
| POST | /workflows/run | Run a hosted workflow and return summary plus artifacts | workflowId + input; respects concurrency limit |
| GET | /workflows/runs/:id | Fetch workflow status and artifact list | User-scoped |
| GET | /workflows/runs/:id/artifacts/:artifactId | Read available workflow artifacts such as evidence JSON, CSVs, Markdown briefs, and reports | Owner scoped; temporary file-backed artifacts can return HTTP 410 |
| POST | /agent/sessions | Open an interactive browser session (stealth + CAPTCHA solving). Returns a session id and watch link | Needs ≥1 credit balance |
| POST | /agent/sessions/:id/screenshot | See the page: screenshot image plus DOM elements with click coordinates and text | Billed per minute of use |
| POST | /agent/sessions/:id/{goto,click,type,scroll,press,read} | Drive the browser — navigate, click, type, scroll, press keys, read the page | 120 credits / minute of use |
| POST | /agent/sessions/:id/replay/{start,stop} | Record an MP4 replay of the session for later review | Multiple per session |
| DELETE | /agent/sessions/:id | Close the browser session and stop the meter | Session-scoped |
| GET | /me | Account info: plan, credit balance, concurrency limit | API key auth |
| GET | /billing/balance | Current credit balance in millicredits | API key auth |
| GET | /billing/credits | Full per-tool cost catalog (live credit prices for every tool) | API key auth |
| GET | /rates | Public pricing contract, including the 15,000-credit/month active-connection add-on and connected function, proxy, and compute rates | 2 Credits per function run, 2 per provider proxy request, and 5 per compute-second, prorated from milliseconds |
| GET | /ledger | Recent credit ledger entries (debits, refunds, top-ups) | User-scoped |
| POST | /api-key/rotate | Rotate your API key | Invalidates the previous key |
| POST | /memory/mcp-call | Generic compatibility gateway to all 126 memory tools — persistent notes, facts, vaults, scheduled actions, tables, channels, research, and CRM | {toolName, args}; 10 MB free / 5 GB Pro storage |
| GET | /memory/connect | Get a ready-to-paste unified MCP connector config for supported clients | Uses MCP_SCRAPER_API_KEY; no separate Memory key is created or returned |

## Limits and safety

- API keys must be sent in the x-api-key header. Query-string keys are rejected.
- /harvest/sync is one continuous request with a 400-second server deadline; maxQuestions is capped at 100.
- /extract-site handles up to 10,000 pages; /map-urls up to 10,000 URLs.
- URL tools only fetch public http/https hosts. Localhost, private IPs, and internal network addresses are blocked.
- Webhook callbacks must be HTTPS public URLs.

## Errors

| HTTP | Meaning | What to do |
|---|---|---|
| 400 | Bad input | Missing field, invalid URL, private URL target, or invalid mode. |
| 401 | Auth | Missing or invalid x-api-key header. |
| 402 | Payment | The requested operation needs available credits. |
| 429 | Limit | User already has an active harvest running. |
| 500/503 | Server or upstream | Inspect error_code and retryable. Retry automatically only when retryable is true; charge_status reports billing, not retry safety. |
