Scrape do with Pydantic AI

How to integrate Scrape do MCP with Pydantic AI

Trusted by

GET STARTED FOR FREE GET A DEMO

30 min · no commitment · see it on your stack

Introduction Also integrate Scrape do with TL;DR What is Pydantic AI What is the Scrape do MCP server Supported Tools & Triggers Creating MCP Server - Stand-alone vs Composio SDK Step-by-step Guide Complete Code Conclusion How to build Scrape do MCP Agent with another framework Explore Other Toolkits FAQ

Connect Scrape do without Auth hassles

We manage OAuth, API Key, token refresh, and scopes, you just build.

Try for Free

Introduction

This guide walks you through connecting Scrape do to Pydantic AI using the Composio tool router. By the end, you'll have a working Scrape do agent that can scrape product prices from a dynamic website, extract news headlines with javascript rendering, bypass cloudflare to get full page html through natural language commands.

This guide will help you understand how to give your Pydantic AI agent real control over a Scrape do account through Composio's Scrape do MCP server.

Before we dive in, let's take a quick look at the key ideas and tools involved.

Also integrate Scrape do with

OpenAI Agents SDK Claude Agent SDK Claude Code Claude Cowork Codex OpenClaw Hermes CLI Google ADK LangChain Vercel AI SDK Mastra AI LlamaIndex CrewAI

TL;DR

Here's what you'll learn:

How to set up your Composio API key and User ID
How to create a Composio Tool Router session for Scrape do
How to attach an MCP Server to a Pydantic AI agent
How to stream responses and maintain chat history
How to build a simple REPL-style chat interface to test your Scrape do workflows

What is Pydantic AI?

Pydantic AI is a Python framework for building AI agents with strong typing and validation. It leverages Pydantic's data validation capabilities to create robust, type-safe AI applications.

Key features include:

Type Safety: Built on Pydantic for automatic data validation
MCP Support: Native support for Model Context Protocol servers
Streaming: Built-in support for streaming responses
Async First: Designed for async/await patterns

What is the Scrape do MCP server, and what's possible with it?

The Scrape do MCP server is an implementation of the Model Context Protocol that connects your AI agent and assistants like Claude, Cursor, etc directly to your Scrape do account. It provides structured and secure access to robust web scraping tools, so your agent can perform actions like scraping dynamic pages, managing sessions, setting custom headers or proxies, and extracting structured data from any website on your behalf.

Dynamic page scraping with headless browsers: Retrieve fully rendered HTML content from JavaScript-heavy or protected websites by leveraging advanced browser emulation and proxy rotation.
Custom scraping session management: Set device type, cookies, wait times, and custom headers to imitate different users, maintain sessions, or access device-specific content for tailored data extraction.
Proxy and anti-bot bypass control: Enable super or proxy modes to utilize residential, mobile, or datacenter proxies, helping your agent bypass strict anti-bot systems and geo-restrictions seamlessly.
Targeted resource filtering: Block specific URLs like ads or analytics scripts during scraping to increase speed, avoid distractions, and improve privacy.
Account usage and statistics retrieval: Access real-time usage stats, subscription status, and remaining request limits so your agent can monitor scraping quotas and avoid interruptions.

Supported Tools & Triggers

Tools

Cancel Async JobTool to cancel an asynchronous scraping job.

Create Async Scraping JobTool to create an asynchronous scraping job with specified targets and options.

Get Account InformationRetrieves account information and usage statistics from Scrape.

Get Amazon Product OffersGet all seller offers for any Amazon product.

Get Amazon product detailsExtract structured product data from Amazon product detail pages (PDP).

Get Amazon raw HTMLTool to get raw HTML from any Amazon page with ZIP code geo-targeting.

Get Async API Account InformationTool to get account information for the Async API including concurrency limits and usage statistics.

Get Async Job DetailsTool to retrieve details and status of a specific asynchronous scraping job.

Get Async Task ResultTool to retrieve the result of a specific task within an asynchronous job.

Scrape webpage using scrape.doA tool to scrape web pages using scrape.

List Asynchronous Scraping JobsTool to list all asynchronous scraping jobs.

Use Scrape.do Proxy ModeThis tool implements the Proxy Mode functionality of scrape.

Scrape URL using POST methodTool to scrape web pages using POST method via scrape.

Search Amazon productsTool to search Amazon and scrape product listings with structured results.

Block specific URLs during scrapingThis tool allows users to block specific URLs during the scraping process.

Set Regional Geolocation for ScrapingThis tool allows users to set a broader geographical targeting by specifying a region code instead of a specific country code.

What is the Composio tool router, and how does it fit here?

What is Composio SDK?

Composio's Composio SDK helps agents find the right tools for a task at runtime. You can plug in multiple toolkits (like Gmail, HubSpot, and GitHub), and the agent will identify the relevant app and action to complete multi-step workflows. This can reduce token usage and improve the reliability of tool calls. Read more here: Getting started with Composio SDK

The tool router generates a secure MCP URL that your agents can access to perform actions.

How the Composio SDK works

The Composio SDK follows a three-phase workflow:

Discovery: Searches for tools matching your task and returns relevant toolkits with their details.
Authentication: Checks for active connections. If missing, creates an auth config and returns a connection URL via Auth Link.
Execution: Executes the action using the authenticated connection.

Step-by-step Guide

Prerequisites

Before starting, make sure you have:

Python 3.9 or higher
A Composio account with an active API key
Basic familiarity with Python and async programming

Getting API Keys for OpenAI and Composio

OpenAI API Key

Go to the OpenAI dashboard and create an API key. You'll need credits to use the models, or you can connect to another model provider.
Keep the API key safe.

Composio API Key

Log in to the Composio dashboard.
Navigate to your API settings and generate a new API key.
Store this key securely as you'll need it for authentication.

Install dependencies

bash

pip install composio pydantic-ai python-dotenv

Install the required libraries.

What's happening:

composio connects your agent to external SaaS tools like Scrape do
pydantic-ai lets you create structured AI agents with tool support
python-dotenv loads your environment variables securely from a .env file

Set up environment variables

bash

COMPOSIO_API_KEY=your_composio_api_key_here
USER_ID=your_user_id_here
OPENAI_API_KEY=your_openai_api_key

Create a .env file in your project root.

What's happening:

COMPOSIO_API_KEY authenticates your agent to Composio's API
USER_ID associates your session with your account for secure tool access
OPENAI_API_KEY to access OpenAI LLMs

Import dependencies

python

import asyncio
import os
from dotenv import load_dotenv
from composio import Composio
from pydantic_ai import Agent
from pydantic_ai.mcp import MCPServerStreamableHTTP

load_dotenv()

What's happening:

We load environment variables and import required modules
Composio manages connections to Scrape do
MCPServerStreamableHTTP connects to the Scrape do MCP server endpoint
Agent from Pydantic AI lets you define and run the AI assistant

Create a Tool Router Session

python

async def main():
    api_key = os.getenv("COMPOSIO_API_KEY")
    user_id = os.getenv("USER_ID")
    if not api_key or not user_id:
        raise RuntimeError("Set COMPOSIO_API_KEY and USER_ID in your environment")

    # Create a Composio Tool Router session for Scrape do
    composio = Composio(api_key=api_key)
    session = composio.create(
        user_id=user_id,
        toolkits=["scrape_do"],
    )
    url = session.mcp.url
    if not url:
        raise ValueError("Composio session did not return an MCP URL")

What's happening:

We're creating a Tool Router session that gives your agent access to Scrape do tools
The create method takes the user ID and specifies which toolkits should be available
The returned session.mcp.url is the MCP server URL that your agent will use

Initialize the Pydantic AI Agent

python

# Attach the MCP server to a Pydantic AI Agent
scrape_do_mcp = MCPServerStreamableHTTP(url, headers={"x-api-key": COMPOSIO_API_KEY})
agent = Agent(
    "openai:gpt-5",
    toolsets=[scrape_do_mcp],
    instructions=(
        "You are a Scrape do assistant. Use Scrape do tools to help users "
        "with their requests. Ask clarifying questions when needed."
    ),
)

What's happening:

The MCP client connects to the Scrape do endpoint
The agent uses GPT-5 to interpret user commands and perform Scrape do operations
The instructions field defines the agent's role and behavior

Build the chat interface

python

# Simple REPL with message history
history = []
print("Chat started! Type 'exit' or 'quit' to end.\n")
print("Try asking the agent to help you with Scrape do.\n")

while True:
    user_input = input("You: ").strip()
    if user_input.lower() in {"exit", "quit", "bye"}:
        print("\nGoodbye!")
        break
    if not user_input:
        continue

    print("\nAgent is thinking...\n", flush=True)

    async with agent.run_stream(user_input, message_history=history) as stream_result:
        collected_text = ""
        async for chunk in stream_result.stream_output():
            text_piece = None
            if isinstance(chunk, str):
                text_piece = chunk
            elif hasattr(chunk, "delta") and isinstance(chunk.delta, str):
                text_piece = chunk.delta
            elif hasattr(chunk, "text"):
                text_piece = chunk.text
            if text_piece:
                collected_text += text_piece
        result = stream_result

    print(f"Agent: {collected_text}\n")
    history = result.all_messages()

What's happening:

The agent reads input from the terminal and streams its response
Scrape do API calls happen automatically under the hood
The model keeps conversation history to maintain context across turns

Run the application

python

if __name__ == "__main__":
    asyncio.run(main())

What's happening:

The asyncio loop launches the agent and keeps it running until you exit

Complete Code

Here's the complete code to get you started with Scrape do and Pydantic AI:

python

import asyncio
import os
from dotenv import load_dotenv
from composio import Composio
from pydantic_ai import Agent
from pydantic_ai.mcp import MCPServerStreamableHTTP

load_dotenv()

async def main():
    api_key = os.getenv("COMPOSIO_API_KEY")
    user_id = os.getenv("USER_ID")
    if not api_key or not user_id:
        raise RuntimeError("Set COMPOSIO_API_KEY and USER_ID in your environment")

    # Create a Composio Tool Router session for Scrape do
    composio = Composio(api_key=api_key)
    session = composio.create(
        user_id=user_id,
        toolkits=["scrape_do"],
    )
    url = session.mcp.url
    if not url:
        raise ValueError("Composio session did not return an MCP URL")

    # Attach the MCP server to a Pydantic AI Agent
    scrape_do_mcp = MCPServerStreamableHTTP(url, headers={"x-api-key": COMPOSIO_API_KEY})
    agent = Agent(
        "openai:gpt-5",
        toolsets=[scrape_do_mcp],
        instructions=(
            "You are a Scrape do assistant. Use Scrape do tools to help users "
            "with their requests. Ask clarifying questions when needed."
        ),
    )

    # Simple REPL with message history
    history = []
    print("Chat started! Type 'exit' or 'quit' to end.\n")
    print("Try asking the agent to help you with Scrape do.\n")

    while True:
        user_input = input("You: ").strip()
        if user_input.lower() in {"exit", "quit", "bye"}:
            print("\nGoodbye!")
            break
        if not user_input:
            continue

        print("\nAgent is thinking...\n", flush=True)

        async with agent.run_stream(user_input, message_history=history) as stream_result:
            collected_text = ""
            async for chunk in stream_result.stream_output():
                text_piece = None
                if isinstance(chunk, str):
                    text_piece = chunk
                elif hasattr(chunk, "delta") and isinstance(chunk.delta, str):
                    text_piece = chunk.delta
                elif hasattr(chunk, "text"):
                    text_piece = chunk.text
                if text_piece:
                    collected_text += text_piece
            result = stream_result

        print(f"Agent: {collected_text}\n")
        history = result.all_messages()

if __name__ == "__main__":
    asyncio.run(main())

Conclusion

You've built a Pydantic AI agent that can interact with Scrape do through Composio's Tool Router. With this setup, your agent can perform real Scrape do actions through natural language. You can extend this further by:

Adding other toolkits like Gmail, HubSpot, or Salesforce
Building a web-based chat interface around this agent
Using multiple MCP endpoints to enable cross-app workflows (for example, Gmail + Scrape do for workflow automation)

This architecture makes your AI agent "agent-native", able to securely use APIs in a unified, composable way without custom integrations.

How to build Scrape do MCP Agent with another framework

OpenAI Agents SDK

Use Scrape do MCP with OpenAI Agents SDK

Claude Agent SDK

Use Scrape do MCP with Claude Agent SDK

Claude Code

Use Scrape do MCP with Claude Code

Claude Cowork

Use Scrape do MCP with Claude Cowork

Codex

Use Scrape do MCP with Codex

OpenClaw

Use Scrape do MCP with OpenClaw

Hermes

Use Scrape do MCP with Hermes

CLI

Use Scrape do MCP with CLI

Google ADK

Use Scrape do MCP with Google ADK

LangChain

Use Scrape do MCP with LangChain

Vercel AI SDK

Use Scrape do MCP with Vercel AI SDK

Mastra AI

Use Scrape do MCP with Mastra AI

LlamaIndex

Use Scrape do MCP with LlamaIndex

CrewAI

Use Scrape do MCP with CrewAI

Explore Other Toolkits

Excel

Oauth2S2s Oauth2

Microsoft Excel is a robust spreadsheet application for organizing, analyzing, and visualizing data. It's the go-to tool for calculations, reporting, and flexible data management.

21risk

Api Key

21RISK is a web app built for easy checklist, audit, and compliance management. It streamlines risk processes so teams can focus on what matters.

Abstract

Api Key

Abstract provides a suite of APIs for automating data validation and enrichment tasks. It helps developers streamline workflows and ensure data quality with minimal effort.

TOOLKIT MARKETPLACE

FAQ

What are the differences in Tool Router MCP and Scrape do MCP?

With a standalone Scrape do MCP server, the agents and LLMs can only access a fixed set of Scrape do tools tied to that server. However, with the Composio Tool Router, agents can dynamically load tools from Scrape do and many other apps based on the task at hand, all through a single MCP endpoint.

Can I use Tool Router MCP with Pydantic AI?

Yes, you can. Pydantic AI fully supports MCP integration. You get structured tool calling, message history handling, and model orchestration while Tool Router takes care of discovering and serving the right Scrape do tools.

Can I manage the permissions and scopes for Scrape do while using Tool Router?

Yes, absolutely. You can configure which Scrape do scopes and actions are allowed when connecting your account to Composio. You can also bring your own OAuth credentials or API configuration so you keep full control over what the agent can do.

How safe is my data with Composio Tool Router?

All sensitive data such as tokens, keys, and configuration is fully encrypted at rest and in transit. Composio is SOC 2 Type 2 compliant and follows strict security practices so your Scrape do data and credentials are handled as safely as possible.

Used by agents from

Never worry about agent reliability

We handle tool reliability, observability, and security so you never have to second-guess an agent action.

Get started for free Get a demo↗

Harsha GaddipatiCo-founder, Slashy

Karan skipped his own birthday party to fix our critical issue. It was 10 pm and he diverted his Waymo to help us instead. This really sets the bar, shows you the commitment you need to have when users rely on your software.

Abhi AryaCo-founder, Opennote

A lot of students tell us that the moment their connected tools start talking to each other inside Opennote feels almost magical. The agent just knows them, and it has immensely helped in keeping new users on the platform.

Nirman DaveCEO, Zams

We chose Composio over Pipedream because it delivered depth where it mattered. It supported niche tools and tricky edge cases that other platforms simply ignored. Giving us confidence to scale without compromising.

Ryan YuFounder, Extra Thursday

As a solo builder, shipping fast is life or death. The only way I can outcompete incumbents is by outmanoeuvring them. Getting bogged down in the complexities of managing agent auth would have been a death sentence for Extra Thursday.

Tomisin JenrolaFounder & CEO, SwarmZero

Before partnering with Composio, adding tool integrations was a slow, resource-intensive process. Each integration could take weeks or months of engineering time, and maintaining them meant constantly keeping up with API changes.

Jerome LeclancheCo-Founder, Ingram Technologies

With hands-on help from their founder, we integrated Gmail and Google Drive in just 30 minutes. This level of personal support and commitment is exactly what startups should strive for.

Harsha GaddipatiCo-founder, Slashy

Abhi AryaCo-founder, Opennote

Nirman DaveCEO, Zams

Ryan YuFounder, Extra Thursday

Tomisin JenrolaFounder & CEO, SwarmZero

Jerome LeclancheCo-Founder, Ingram Technologies

With hands-on help from their founder, we integrated Gmail and Google Drive in just 30 minutes. This level of personal support and commitment is exactly what startups should strive for.

Harsha GaddipatiCo-founder, Slashy

Abhi AryaCo-founder, Opennote

Nirman DaveCEO, Zams

Ryan YuFounder, Extra Thursday

Tomisin JenrolaFounder & CEO, SwarmZero

Jerome LeclancheCo-Founder, Ingram Technologies

With hands-on help from their founder, we integrated Gmail and Google Drive in just 30 minutes. This level of personal support and commitment is exactly what startups should strive for.

How to integrate Scrape do MCP with Pydantic AI

Table of Contents

Connect Scrape do without Auth hassles

Introduction

Also integrate Scrape do with

TL;DR

What is Pydantic AI?

What is the Scrape do MCP server, and what's possible with it?

Supported Tools & Triggers

What is the Composio tool router, and how does it fit here?

What is Composio SDK?

How the Composio SDK works

Step-by-step Guide

Prerequisites

Getting API Keys for OpenAI and Composio

Install dependencies

Set up environment variables

Import dependencies

Create a Tool Router Session

Initialize the Pydantic AI Agent

Build the chat interface

Run the application

Complete Code

Conclusion

How to build Scrape do MCP Agent with another framework

OpenAI Agents SDK

Claude Agent SDK

Claude Code

Claude Cowork

Codex

OpenClaw

Hermes

CLI

Google ADK

LangChain

Vercel AI SDK

Mastra AI

LlamaIndex

CrewAI

Explore Other Toolkits

Excel

21risk

Abstract

FAQ

What are the differences in Tool Router MCP and Scrape do MCP?

Can I use Tool Router MCP with Pydantic AI?

Can I manage the permissions and scopes for Scrape do while using Tool Router?

How safe is my data with Composio Tool Router?

Used by agents from

Never worry about agent reliability