← Back to blog

Company Data API Guide: Endpoints, Schemas, and Examples

You're usually staring at the same problem from three angles at once. Sales wants funding alerts before a competitor gets there, recruiting wants fresh company signals before the job board goes stale, and an agent builder wants structured records the model can call without brittle scraping glue. A company data API is the piece that turns those needs into something programmable, but the decision isn't just “which vendor has the biggest database.” It's whether the feed is fresh enough, structured enough, and delivered in the right way for the workflow you're trying to automate.

Table of Contents

What a Company Data API Does

A company data API gives software a direct way to request business records, without the manual work of research and cleanup. An SDR can look up a newly funded startup by domain, a recruiter can pull leadership contacts, and an analyst can query firm history through one interface instead of downloading files and reconciling them by hand. The value is repeatable access to the same schema every time.

That consistency matters because public and commercial business data now ships as structured endpoints, not only as files. The U.S. Census Bureau's business-data API catalog includes long-running datasets such as Business Dynamics Statistics from 1978–2023, Economic Census from 2002–2022, and County Business Patterns from 1986–2021, with fields like establishments, payroll, employment, revenues, and assets exposed through programmatic access rather than manual downloads Census Bureau business data API catalog. The Census Bureau also says its API supports custom queries and embedding statistics into web or mobile apps, which is why APIs have become the default delivery layer for company and market data.

Practical rule: if a vendor cannot tell you what changes are queryable, and how fast those changes surface, the dataset will drift out of sync with your workflow.

A production-grade company data API usually accepts identifiers like domain, legal name, or registry number, then returns normalized JSON with firmographic fields, funding rounds, hiring signals, and verified contacts. Better systems also expose confidence metadata, so automation can decide whether to route a lead, queue it for review, or ignore it.

The delivery path matters as much as the record itself. For example, NowFunded is the kind of source buyers evaluate on freshness and contact quality, because stale funding data or weak verification logic breaks outbound workflows fast.

The key shift is that the API is not just a lookup tool. It becomes the distribution layer for SDR routing, recruiting sourcing, market intelligence, and agent tool calls. When that layer is well designed, the same record can power a human dashboard, a webhook, and an MCP tool without forcing you to rebuild the data contract each time.

Delivery Models Compared

The market mostly comes down to four delivery shapes, and each one solves a different job. REST is the workhorse for direct lookups and enrichment. Webhooks are for event delivery when freshness matters more than convenience. MCP servers are the cleanest fit when the consumer is an AI agent. CSV or dashboard exports are useful for offline analysis, but they're the weakest choice for operational workflows.

A REST call is what most buyers expect first. A request like GET /v1/companies?domain=... returns a JSON record you can parse immediately, which makes it ideal for form fill enrichment, CRM augmentation, and one-off searches. The trade-off is obvious, you still have to ask for change, so REST pushes you toward polling unless the provider offers another channel.

Webhooks flip that model. The provider sends a POST to your endpoint when a funding event, executive change, or status update occurs, which is the right shape when the business value depends on speed. That's why webhook delivery is the strongest fit for outbound teams and real-time alerting, provided your receiver is built to handle retries and deduplication.

MCP delivery is newer in this space, but it fits the agent-native use case well. Instead of wrapping APIs in custom orchestration glue, you expose tools such as search and contact lookup directly to the model. That lowers integration friction for builders, but it does mean you're now hosting and securing a tool server, not just a web endpoint.

CSV exports still have a place for backfills and spreadsheet-heavy analysis, but they're poor as a primary feed. The moment a schema changes or a row arrives late, your downstream pipeline inherits the breakage. Use exports for batch work, not as the core of a live system.

Delivery Model Best For Latency Operational Burden REST Form fill enrichment, one-off lookups, CRM sync On demand Low to moderate Webhook Funding alerts, executive changes, time-sensitive triggers Near real time Moderate to high MCP AI agents, tool-based workflows, conversational retrieval On demand per tool call Moderate CSV or dashboard export Backfills, analyst slicing, offline review Slow, batch-based Low at first, high later

The buying question is rarely “which model exists.” It's “which model matches the job without forcing my team into polling loops or manual exports.” If the answer changes by persona, the API is probably healthy. If the answer is one-size-fits-all, the product is usually hiding trade-offs instead of solving them.

JSON Schema Anatomy for a Funding Record

A funding record is where schema quality becomes visible. If the provider exposes sloppy fields here, the rest of the platform usually follows the same pattern. A serious implementation should treat funding as a typed object, not a loose blob of text.

The fields that need to be exact

Start with the root structure. The record should wrap an array of funding rounds, and each round should carry a required round_id so downstream systems can insert idempotently. That matters because webhook retries, scheduled syncs, and manual replays all need a stable key for deduplication.

The announced_date field should use a proper date format, not an untyped string. If you're joining it to press releases, CRM notes, or analyst workflows, malformed dates create silent mismatches. amount_raised_usd should be a number with a minimum of zero, and if the provider can't disclose it, the schema should allow null rather than forcing consumers to guess from text.

round_type should be an enum, not a free-text field. Once you let buyers write their own synonyms, analytics collapses under variations like Series A, series-a, Series-A, and other near-duplicates. Investors should be an array of objects with a name and lead boolean, because the lead investor often matters more than the full syndicate for routing and messaging.

A clean schema doesn't make the data more truthful by itself, but it does make bad data easier to detect before it hits your CRM.

Confidence belongs in the payload too. A 0 to 1 confidence field gives the caller a way to decide whether to auto-route, hold for review, or enrich again later. That's the difference between a feed that supports automation and one that only supports browsing.

A pattern that works in production

The most useful schema pattern is simple, explicit, and boring. Required fields should be the ones your workflow can't live without, usually domain, round_id, announced_date, and round_type. Everything else can be optional, but it still needs documentation and a stable type. In OpenAPI 3.1, Schema Objects align with JSON Schema Draft 2020-12, which makes validation, typed client generation, and contract testing much more consistent Manning, The Design of Web APIs, chapter 7.

That alignment is what lets automation trust the feed. If the schema says email is a string with a proper format, downstream tools can validate before sending. If HQ location is nested, clients don't have to infer whether it belongs to the company or the contact. For company intelligence, ambiguity is expensive.

REST Endpoint Patterns

A production REST API should read cleanly from the path. A versioned base path like /v1/companies signals that the contract can change without breaking clients, which matters when sales, recruiting, and automation teams all depend on the same dataset.

Search and detail calls do different jobs

Search endpoints are for discovery. A call like GET /v1/companies/search?name=...&industry=... should return ranked results, match scores, and light fields that help a user or UI pick the right record fast. Detail endpoints are for retrieval. A call like GET /v1/companies/{id} or GET /v1/companies/{domain} should return the full object, including funding history, employee counts, social profiles, and verified contacts.

The response shape should match the job. Search results need concise snippets and filtering metadata. Detail responses need depth. Mixing the two creates heavy search pages and thin detail pages, which slows integration work.

Dimension Search Endpoint Detail Endpoint Primary purpose Find the right company Retrieve the full company record Typical response Ranked list, match score, light fields One canonical record, deeper nested objects Best user SDR, analyst, UI browsing flow CRM sync, enrichment, automation Risk if misused Expensive broad scans Overfetching when only a match is needed

Pagination deserves real attention. Cursor-based pagination usually works better than offset pagination for large company datasets because it holds up when records are added or updated between requests. Offset-based paging is simpler on paper, but it drifts as the dataset changes, which is a poor fit for live business data. Server-side filter whitelisting matters too, because letting buyers combine every field with every operator can create expensive queries that do not scale.

Authentication should be obvious in the contract. The consumer should know where the key goes, what is scoped, and whether the response includes rate-limit or usage headers. If a vendor makes you guess, the integration will be fragile. For schema design details that affect payload validation, see the next section on JSON Schema.

Polling, Webhooks, and MCP Servers

A company data API lives or dies on delivery mechanics. The questions are whether you can keep records fresh without hammering the service, whether your receiver can survive retries and duplicates, and whether the schema stays stable enough for downstream systems to trust it.

Polling is still the easiest path to ship, and it remains common for backfills and batch enrichment. A client calls a GET endpoint on a schedule, then uses ETag or If-Modified-Since to skip unchanged records. The trade-off is freshness. Short intervals raise load and cost, while long intervals leave SDR workflows and recruiting queues working from stale data.

Webhook delivery shifts the burden to the receiver. The provider POSTs a change event, your app verifies the signature, checks the payload against the expected shape, and stores the event before business logic runs. That gives you faster alerts for funding or contact updates, but it also means you need durable retry handling, deduplication by event_id, and a clear response plan for temporary downtime.

MCP adds a different operational model for agent builders. Instead of wrapping REST calls in custom orchestration, you expose tools such as search_company or get_funding_rounds, then control access, logging, and schema validation at the tool boundary. The upside is less glue code. The burden shifts to tool governance, rate limits, and monitoring each invocation path.

Criterion Polling Webhook MCP Server Retry handling Simple on the client side Must be designed into the receiver Managed per tool call Duplicate control Usually handled by the poller Required at ingress Required at the tool layer Infrastructure burden Low Moderate to high Moderate Operational risk Stale reads Missed events if the receiver fails Tool sprawl and access control

For operators, the checklist is straightforward. Use polling when the system can tolerate scheduled freshness and simple recovery. Use webhooks only if you can verify signatures, persist incoming events immediately, and dedupe retries without manual cleanup. Use MCP when the consumer is an agent and the schema is tight enough that every tool call can be validated, logged, and governed.

Webhook Reliability Checklist

Webhook delivery breaks in predictable ways. A provider retries, a receiver processes the same event twice, or the payload lands while the app is down. The fix is a disciplined receiver, not extra logging after the fact.

A checklist of six best practices for ensuring the reliability and security of webhook implementations.

The checks that keep data from disappearing

Start with a unique event_id on every payload. Store it before business logic runs, then reject duplicates on receipt. That field handles most accidental double-processing when retries happen.

Verify the signature before you parse the body. A signed webhook with timestamp checks blocks forged payloads and replay attacks, which matters any time funding alerts or contact updates can trigger outbound actions. Keep the raw payload intact in durable storage first, then process it asynchronously.

Retries need a clear policy. The provider should use backoff, stop after a defined ceiling, and send failures to dead-letter handling if the receiver keeps failing. Your side should return a fast 2xx acknowledgment and move heavy work downstream so the sender does not keep hammering your endpoint.

  • Dedupe by event ID: Store processed IDs for a meaningful retention window so a replay does not create a second CRM update.
  • Validate the signature first: Reject anything unsigned or stale before it touches your parser.
  • Persist raw payloads immediately: Keep the original event so you can replay or audit later.
  • Process asynchronously: Do not let webhook delivery block on enrichment, routing, or Slack notifications.
  • Expose replay support: If the receiver goes down, a manual resend path saves hours of cleanup.
  • Alert on exhaustion: Dead-letter queues and retry failure should page the right person fast.

A strong webhook pipeline feels uneventful because the rough edges are already handled. That is the goal. If every update turns into a support ticket, the delivery model is wrong or the receiver is too fragile.

Verified Contacts and Pay-on-Success Pricing

Contact data changes the economics of a company data API. A company record without a usable contact can still help with research, but outbound teams live and die by whether the address works. If the billing model ignores that reality, the buyer pays for noise.

A comparison chart showing that a pay-on-success model saves money compared to a flat subscription for data verification.

Why the billing model matters

Flat subscription pricing is easy to understand, but it can hide a lot of waste. You pay the same whether the contact is deliverable, role-based, risky, or unusable. Pay-on-success changes the incentive, because the vendor only bills when verification or delivery succeeds.

The useful tiers are usually behavioral, not just descriptive. Safe-to-send contacts can go straight into outreach. Role-based inboxes need caution. Catch-all domains require more review. Risky and unknown records should usually stay out of the first pass. Those categories belong in the schema as confidence or verification metadata, not buried in a dashboard nobody opens.

Operational rule: if your campaign depends on reply rates, buy verified contacts, not raw pulls.

That distinction matters at the unit-economics level. The brief's example shows how a pay-on-success model can change a 10,000-lookup campaign from a flat $500/month pattern to a pay-for-success path that lands at $178.50 in the illustrated scenario, based on verification attempts, verified records, and successful deliveries. Those are vendor-specific pricing structures, not universal market rates, but they show why the billing model can change whether a campaign is viable.

The best APIs expose that logic clearly in the response. If the verified flag, deliverability_status, and confidence fields all line up, an SDR can route the record without a manual cleanup layer. If they don't, the buyer ends up paying twice, once for the pull and again for the missed opportunity.

Worked Examples for Two Real Users

Morgan is running outbound for a startup team, and speed is everything. She subscribes to a funding-event webhook, receives a POST when a Series B is verified, enriches the company with a contact lookup, and pushes the result into her CRM with a Slack ping. The important thing isn't the mechanics, it's the sequence, because the webhook removes the polling delay that usually costs the first touch.

A diagram illustrating how two different user roles use a company data API to improve business operations.

Morgan's flow

Her receiver gets a signed payload that includes the company identifier, round type, and announced date. The app validates the event, stores it, and then makes a follow-up request for verified founder contacts. From there, the CRM update happens automatically, and the SDR sequence launches while the signal is still fresh.

Dev's flow

Dev is building an indie agent workflow, so he exposes the same backend through an MCP server. His agent can call search_company, get_funding, and get_verified_contacts without custom orchestration code, then answer queries like which SaaS startups raised recently and have verified founder emails. He tolerates a little more per-call overhead than Morgan does, because the model is the consumer and the workflow is conversational.

You can see the difference in the auth shape too. Morgan's webhook receiver authenticates the incoming POST. Dev's MCP server authenticates tool access. The underlying JSON schema stays the same, which is what keeps the stack from fragmenting as the consumer changes.

NowFunded's blog is a useful reference point for how a funding feed can be packaged for both human and machine use, but the broader lesson is vendor-agnostic. If the data contract works for push and pull, you don't need a separate product for every persona.

Cross-Market Consistency and Entity Resolution

A global company dataset sounds clean until you start matching records in production. The same legal entity can appear differently in Delaware filings, UK Companies House, LinkedIn, and commercial databases, each with its own identifier and update cadence. If a vendor claims global coverage without explaining how those records are reconciled, the buyer is inheriting hidden risk.

Entity resolution is the process of collapsing those signals into one canonical company_id. The strongest matching signals are the ones that survive across systems, like domain, normalized legal name, jurisdiction-specific registration number, employee count range, and HQ geography. Vanity signals such as logo color or tagline create false positives, which is how bad merges slip into outbound and recruiting workflows.

Jurisdiction details matter more than most marketing pages admit. French records can rely on different identifiers than German ones, and Israeli formats won't always line up with U.S. expectations. That's why a company data API should surface jurisdiction-aware confidence or matching metadata instead of pretending every market fits one clean schema.

If the record looks universal but the match logic is invisible, the dataset is probably doing more guessing than the buyer realizes.

The practical test is simple. Ask how the provider decides two records are the same company, what fields are authoritative in each market, and how confidence changes when sources disagree. If the answer is vague, the integration will spend its life correcting merges instead of using the data.

Buyer Evaluation Checklist

A good evaluation doesn't need a committee. It needs a short scorecard that punishes vague claims and rewards operational clarity. Thirty minutes is enough to tell whether a vendor is ready for production or just polished for demos.

A buyer evaluation checklist for company data APIs featuring five criteria for assessing data quality and services.

Score the API on the things that matter

Start with schema documentation quality. Look for field-level examples, explicit required properties, and a JSON Schema URL or equivalent contract reference. If you can't tell how a field should be parsed, the API isn't production-ready.

Then check freshness signals. Ask when the last funding date was updated, how recent a hire or executive change is, and whether the vendor can show verified update timestamps. After that, review delivery options. A serious platform should explain REST, webhooks, and batch export clearly, with no guessing about where each one fits.

Finally, evaluate the billing model and coverage for your target segments. Pay-per-success contact pricing should be transparent, and the vendor should tell you how it treats verified versus unverified records. If the answers keep drifting into marketing language, that's your signal to move on.

  • Schema clarity: Are fields documented with examples and types?
  • Freshness proof: Can the vendor show recent update timing?
  • Delivery fit: Do REST, webhook, and batch options match your workflow?
  • Contact economics: Is pricing tied to verification success or just raw pulls?
  • Coverage realism: Do the target segments match your use case?

Run the same query three times across 48 hours. If the answer changes in ways the vendor can't explain, freshness claims are probably weaker than the brochure suggests.

Quick-Reference Table by Persona

Different users buy the same API for different reasons, and that should change what you validate first. The table below is the fastest way to map the delivery model to the job.

Persona Best Delivery Model Priority Fields Acceptable Freshness Outbound SDR Webhook, then REST fallback Funding round, lead investor, verified founder contact As close to real time as possible Recruiter REST with filterable search Headcount, HQ, leadership contacts, industry tag Recent enough for active sourcing Market analyst REST plus CSV export for backfill Historical funding, revenues, employment, assets Periodic refresh with stable history AI agent builder MCP server Search, funding, verified contact status On demand per tool call

For SDRs, freshness and verified contacts dominate. For recruiters, structure and searchability matter more. For analysts, historical depth and stable exports are the priority. For agent builders, the schema has to be callable, typed, and easy for models to consume.

FAQ on Company Data API Buying

How is pricing usually structured? It's commonly per record, per seat, per usage tier, or pay-on-success for verified contacts. If the vendor can't explain which actions trigger charges, the model will be hard to forecast.

What freshness SLA is realistic for funding and contact data? Funding data often lags the underlying event because verification takes time, while contact freshness depends on whether the provider is checking deliverability or just returning a stored record. Ask for update timing in plain language, not marketing adjectives.

Can I get a trial or sandbox key without a sales call? Many vendors offer some kind of preview or limited-access plan, but the useful question is whether that trial includes the fields you need. A sandbox that hides contact data doesn't tell you much about production fit.

How long does onboarding take for webhooks or MCP? A documented schema and a clear auth model should let a capable team wire in the basics quickly, but the timeline depends on how much receiver logic, replay handling, and agent orchestration you need. The simpler the payload contract, the faster the integration.

If you want a serious buying pass, use the evaluation checklist above, then test the same query more than once and compare the payload shape, freshness, and delivery behavior. That's the quickest way to see whether a vendor is built for production or just for demos.


If you're building funding alerts, verified contact workflows, or agent-ready company intelligence, NowFunded gives you a live feed designed for exactly that kind of delivery problem. Visit NowFunded to see how a structured company data API can replace stale lists, brittle scraping, and manual cleanup in your stack.