Skip to main content
Back to Blog

Data Legion Is Now on Firecrawl Alexandria

October 6, 2026 · 5 min read

An agent that researches a company usually needs two kinds of data. It needs the open web: the company's site and what's been written about it. And it needs structured records the web won't hand over cleanly, such as who works there, in what role, and how the team is changing. Until now those came from separate tools with separate keys.

Firecrawl's answer is Alexandria, which puts the live web, official data providers, and Firecrawl's own indexes behind one connection. Data Legion's person and company enrichment is now part of it. If your agent already calls Firecrawl, it can turn an email or a LinkedIn URL into a person record, or a domain into a company record, without a second integration.

This post covers what the listing includes, the call itself in Python and JavaScript, the one setup step before the first request, and when the direct Data Legion API is the better fit.

What Alexandria is

Alexandria is a catalog of providers and capabilities that agents reach through Firecrawl's existing API. Each capability has defined inputs, a response contract, and a listed price, and Firecrawl handles discovery, terms, and billing across all of them. Firecrawl launched it alongside its Series B, and reports that agents using Alexandria scored 21 percent higher on answer quality than the same agents using built-in web tools in its internal evaluations.

For an agent builder, the practical change is that a data provider becomes one more call shape on a client you already have. You don't sign up with the provider or keep a second key.

What's in the Data Legion listing

Data Legion appears under People and Company with six capabilities, one per data product:

Capability Returns
people/enrich-base-no-contact Identity, current role and company, work history, education, skills, social profiles
people/enrich-base-with-contact The base record plus emails, phones, location, age, and sex
people/enrich-premium-no-contact The base record plus seniority, job function, decision-maker flag, tenure, and data freshness
people/enrich-premium-with-contact The premium record plus emails, phones, location, age, and sex
companies/enrich-base Firmographics: name, domains, industry, size, type, founding year, description
companies/enrich-premium Firmographics plus headcount, growth, turnover, and workforce breakdowns

Person capabilities take an email, an email hash, a phone number, a social profile URL, or a name paired with one more detail such as a company or city. Company capabilities take a domain, a name, a LinkedIn ID, or a ticker. The response is the same one the Person Enrich and Company Enrich endpoints return: a matches array with the record and its match_metadata, plus total.

Each capability is billed in Firecrawl credits, and only when it returns a match. The current credit cost of each one is shown on the listing in your Firecrawl dashboard.

Calling it

An Alexandria call is a Firecrawl scrape with an alexandria block instead of a URL. Name the provider, the capability, and the options, which are the same fields you'd send to the Data Legion API.

import os

from firecrawl import Firecrawl

firecrawl = Firecrawl(api_key=os.environ["FIRECRAWL_API_KEY"])

result = firecrawl.scrape(
    alexandria={
        "provider": "datalegion",
        "capability": "people/enrich-premium-no-contact",
        "options": {"social_url": "https://www.linkedin.com/in/janedoe"},
    },
)

data = result.alexandria[0].data
person = data["matches"][0]["person"] if data.get("matches") else None
if person:
    print(person["full_name"], person["job_title"], person["seniority_level"])
import Firecrawl from "@mendable/firecrawl-js";

const firecrawl = new Firecrawl({ apiKey: process.env.FIRECRAWL_API_KEY });

const result = await firecrawl.scrape({
  alexandria: {
    provider: "datalegion",
    capability: "people/enrich-premium-no-contact",
    options: { social_url: "https://www.linkedin.com/in/janedoe" },
  },
});

const data = result.alexandria[0].data;
const person = data?.matches?.[0]?.person ?? null;
if (person) {
  console.log(person.full_name, person.job_title, person.seniority_level);
}

Company enrichment is the same call with a different capability and a domain:

result = firecrawl.scrape(
    alexandria={
        "provider": "datalegion",
        "capability": "companies/enrich-base",
        "options": {"domain": "hubspot.com"},
    },
)
company = result.alexandria[0].data["matches"][0]["company"]
const companyResult = await firecrawl.scrape({
  alexandria: {
    provider: "datalegion",
    capability: "companies/enrich-base",
    options: { domain: "hubspot.com" },
  },
});
const company = companyResult.alexandria[0].data.matches[0].company;

Each entry in alexandria also carries the credit cost of that call (creditsCost in JavaScript, credits_cost in Python), which makes per-agent cost tracking a matter of summing one field. A failed call comes back with an error object instead of data, so check for it before reading the record.

One step before the first call

Alexandria asks each organization to accept a provider's data terms before using it. Until an admin does, the request fails with a 403 and the code THIRD_PARTY_DATA_TERMS_REQUIRED, which both SDKs raise as an exception (ProviderTermsRequiredError in Python). The error carries a requiresAction.url where an admin reviews and accepts the terms. Firecrawl's own guidance is that agents need explicit user authorization before accepting terms, so plan for a person to approve this once, in the dashboard, rather than having the agent click through it.

Letting the agent find it

You don't have to hardcode the capability. Alexandria exposes a tool search: Firecrawl's find_tools (or findTools in JavaScript), or a search request with sources: ["alexandria"], returns matching providers with their capability paths, inputs, response contracts, and pricing. Discovery is free. An agent that asks for "work email for a person" or "company headcount" can find the Data Legion capability and call it with the right options.

The usual care with agent tool calls still applies. A match can come back with a field your workflow needs left empty, and the agent won't notice unless you check. Evaluating Agent Tool Calls walks through fill-rate and confidence assertions that work the same on an Alexandria response, since the record inside is the standard Data Legion one.

Alexandria or the API directly

The two routes reach the same data, so the choice comes down to where your agent already lives.

Alexandria fits when the agent runs on Firecrawl. One key, one bill, and enrichment sits next to search and scrape in the same client. It also suits agents that pick their own tools at runtime, because discovery is built in.

The Data Legion API fits when you need more than enrichment or more control. Search and natural-language discovery over the full dataset, bulk delivery, higher rate limits, and enterprise terms all live there. So do our own SDKs, MCP server, and CLI, which a coding agent like the one in A Prospect Research Agent with Claude Code can use directly.

Many teams will use both: Alexandria inside the research agent, and the API in the pipelines that enrich records as they arrive.

To try it, open Alexandria in your Firecrawl dashboard, find Data Legion under People, accept the terms, and run a capability from the playground. The field reference for every capability is in the person and company schemas.

See the data for yourself

Start a free trial and enrich your own records in minutes, or book a demo to talk through enterprise needs.