Blog/Product
11 min read

AI Agent Development Company or In House? The Hard Part

Two vendors, one domain, two different companies. One reported 2,930 employees and the other 63. Neither returned an error. That is the build.

AI Agent Development Company or In House? The Hard Part

Three vendors sell the same thing: give us a domain, we return the company. We called three of them with figma.com and compared what came back.

input: figma.com

hunterio             72 leaf fields    "Figma"          2,930 employees
contactout           24 leaf fields    "Figma Weave"       63 employees
the-companies-api    a 502 error, formatted as JSON

Two of them answered, neither errored, and they returned different companies.

That is the honest version of the build versus buy question. Not whether your team can write an HTTP request, which they obviously can, but who owns the afternoon where somebody works out which of those two rows is the company you meant.

Fair disclosure. You are on the Monid blog, Monid sells access to all three of those endpoints, and the consultancy named near the end is a content partner. The section before it says when none of this needs outside help.

How hard is the API call itself?

Not hard at all, and pretending otherwise is how these arguments get lost.

The part that genuinely is an afternoon

Getting a key, reading a schema, making an authenticated request and parsing the response is a solved problem, and any competent developer does it before lunch. If somebody quotes you a project to call one documented endpoint, the quote is not for the call.

This matters because the usual pitch for outside help leans on complexity that is not there. Agent frameworks are well documented, model providers ship SDKs, and the tool-calling loop is a few dozen lines. A team that can ship a web app can ship the first version of an agent, and they should.

Where the estimate goes wrong

The estimate breaks at the second vendor, not the first.

One endpoint is an afternoon. Two endpoints that answer the same question in different shapes is a mapping layer, a decision about which to trust, and a test that catches the day one of them changes. Five is a small internal product with an owner, and nobody scoped that at the start because the first call took an afternoon and the pattern seemed obvious.

What happens when two vendors answer the same question?

They disagree, and the disagreement is not in the format. It is in the answer.

Same domain, different company

We sent figma.com to two providers that both advertise company enrichment from a domain:

                 hunterio              contactout

name             Figma                 Figma Weave
employees        2,930                 63
linkedin         company/figma         company/figmaweave
founded          2012                  absent
industry         Professional Svcs     Technology, Information and Media

Read the employee row twice. 2,930 against 63, a factor of forty six, and it is not a data quality problem in the usual sense. Both numbers are probably right. They describe different entities: the company and a product line inside it that has its own brand, its own LinkedIn page, and its own subdomain.

Neither response contained a warning. Neither returned an error. An agent routing on company size sends one of these to the enterprise flow and the other to self serve, and it will do so confidently.

The shapes have almost nothing in common

Beyond the values, the two responses barely share a vocabulary. Flattening both to leaf field names:

hunterio           72 leaf fields
contactout         24 leaf fields
shared names        8   of 79 distinct   Jaccard 0.10

The eight in common are name, domain, industry, description, country, followers, type, employees. Everything else is unique to one vendor.

And the nesting is unrelated. Headcount from one is metrics.employeesCount. From the other it is companies["figma.com"].employees, because that provider returns a map keyed by the domain you sent it. Those are not two dialects of one format. They are two formats, and the code that reads one cannot read the other.

That is the part nobody puts in the estimate: swapping a vendor is not a configuration change, it is a rewrite of every accessor that touched the old shape.

A shared name is not a shared meaning

Even the eight names in common are less reassuring than they look. type appears in both responses and means something different in each. followers is a social count in one and a follower total across unspecified networks in the other. employees is a number in one and, in the other, sits inside the domain-keyed map with no indication of what it counted or when.

So the real overlap is smaller than eight. A mapping written by matching field names, which is the obvious first approach and the one an LLM will suggest, produces a table where two columns quietly hold different quantities under one header. The same problem shows up when comparing firmographic providers on cost: the cheaper record is only cheaper if it answers the same question.

Why does a failed call look like a successful one?

Because the failure arrives with the same content type, the same structure, and enough fields to look like data.

What the third provider returned

The third endpoint gave back this, and nothing else:

title        Error 502: Bad gateway
status       502
error_name   origin_bad_gateway
retryable    true
zone         api.parse.bot
ray_id       a3830798ab32ad66

Seventeen fields, valid JSON, a type key at the top just like a real record has. A pipeline that checks whether it got an object back, rather than which object, will write this into a row and move on. The field named type even collides with a legitimate field name in the other two responses.

This is a different failure from a malformed request, which a wrong enum returning zero covers: there the request was wrong and the response was well formed and empty. Here the request was fine and the transport failed, and the response is well formed and full of the wrong thing.

The rule that catches both

Assert on a field you actually need before you accept a response. Not is this an object, not did it return 200, but does it contain a company name. That check is two lines and it is the difference between a pipeline that fails loudly and one that fills a table with error envelopes for a week.

Which of this does a team have to own?

Four jobs, and only the first is the one people budget for.

Calling and mapping

The call, then a normalisation layer that turns every vendor shape into one internal shape. Write it once and every later vendor is a small adapter. Skip it and vendor names leak into your business logic, which is the change that gets expensive later.

Deciding who to believe

The genuinely hard one, because it needs a rule rather than code. When two sources disagree, which wins, and on which field? Freshest? Highest coverage? Vendor A for headcount and vendor B for industry? There is no general answer, only one that fits your use, and somebody has to write it down.

Worth noticing that the Figma case is not a tie to break. One vendor resolved the domain to the parent company and the other to a sub-brand, and both were defensible readings of a request that never specified which was wanted. A precedence rule does not fix that, because the rule assumes the two rows describe the same entity. What fixes it is deciding, at the point the question is asked, what "the company at this domain" means for your product, then testing whether each vendor agrees with that definition. That decision is a product decision wearing an engineering hat, which is why it tends to sit unowned.

Noticing when it changes

Vendors reshape responses, add fields and re-resolve entities without telling you. The test that catches it is not a unit test, because tool responses drift on their own and the world moves under the assertion. What works is checking a fixed set of known records on a schedule and alerting on the diff.

Paying for it in the right shape

Enrichment runs in bursts, so the bill should too. Metered per call fits that better than a seat you hold through the quiet months, which is the argument in paying per call rather than a subscription.

Reaching several vendors through one key removes the procurement half of this, though not the reconciliation half. Monid is the OpenRouter for agent tools: one balance across the enrichment endpoints and everything next to them, with discovery and schema inspection free, so comparing two response shapes costs nothing before you commit. That is the tools half of models and tools being separate integrations, and current per-endpoint figures live at monid.ai/tools rather than in a sentence that will age.

Give this to your agent

$Set up https://monid.ai/SKILL.md, and then use Monid to call two company enrichment endpoints with the same domain and show me every field where they disagree.

When should you hand the build to someone else?

Three cases, and the first two are the common ones.

When the reconciliation is the product

If your agent's value is the judgement about which answer is right, that judgement is yours and should not be outsourced. Buy the calls, own the rules. A team that hands out the part that differentiates them ends up with a system nobody internally can reason about.

When the deadline is shorter than the learning curve

The four jobs above are learnable and take a quarter or so to learn properly, mostly by getting them wrong. If you have that quarter, spend it: the knowledge stays. If you have six weeks and a commitment, buying the experience is cheaper than buying it back later, and this is the honest case for an outside team.

The version of this that goes badly is hiring out the build and keeping none of the reasoning. You get a working system and a set of precedence rules nobody internally can defend, and the first time a vendor re-resolves an entity there is no one who knows why the rule said what it said. If you do bring people in, the deliverable to insist on is the written decisions, not just the repository. The narrower version of the same trade, one job rather than a whole system, is laid out in buy or build email enrichment.

When it is one step inside something larger

Sometimes the data layer is a small part of a system nobody on the team has built before, and the sensible move is to bring in people who have. Consultancies such as Winder.AI work on exactly this end of it, the harness around a model rather than the model, which is a different purchase from an API and occasionally the right one. The equivalent decision for scraping specifically is worked through in should you hire that out, and the same three questions apply here.

The tell that you are in the wrong case: a build that has been two weeks from done for two months, where every week's blocker is a different vendor's response shape. That is not a resourcing problem you can hire your way out of quickly, and it is worth naming before the fourth vendor is added.

Conclusion

One domain, three vendors, and three numbers worth keeping.

Eight shared field names out of seventy nine. 2,930 employees against 63, for the same input, from two providers that both returned successfully. And a 502 error formatted as a seventeen field JSON object that a careless pipeline will store as a company.

None of that is visible when you read the pricing pages, and none of it is hard to fix once you know it is there. The API call is an afternoon. Knowing which answer is true is the build, and whether you do that in house depends on whether that judgement is your product or somebody else's problem.

FAQ

Does using one vendor avoid this?

It hides it rather than avoids it. A single vendor is internally consistent, so the disagreement disappears from view along with the information that there was one. You still get the wrong entity sometimes, you just have nothing to compare it against. Single vendor is a reasonable choice for a first version; treating its answer as ground truth is the mistake.

Two providers disagree. How do I pick?

Per field rather than per vendor. Take a sample of records you can verify by hand, score each provider field by field against them, and write the winners down as a rule. That exercise takes a day and it is the artefact that outlives whichever vendor you started with. Picking a favourite provider and trusting it everywhere is how the Figma versus Figma Weave case gets into production.

Is this the same question as hiring a scraping service?

Related, and priced differently. A managed scraping service is priced against the value of the data rather than the cost of the requests, which is a separate calculation, and it is worked through properly in the managed scraping pricing breakdown.

What does the comparison itself cost to run?

Almost nothing, which is the point. Reading two schemas is free, and calling both endpoints once with the same input is fractions of a cent. The expensive version is discovering the disagreement six months later inside a customer-facing number. Run the comparison first, on twenty of your own records, before the architecture assumes an answer.

Last updated September 2026.

ai agentsbuild vs buydata qualityintegration