The fastest way to work out how to choose an AI CRM is to run a 60-minute architecture test before you sit through a single feature demo. Split it into three 20-minute blocks: the data model, whether the agents run work or only suggest it, and who owns your context and your AI. At the end you have a clean yes or no on whether the vendor is genuinely AI-native, and whether it earns the weeks of procurement that follow.

Most AI-CRM evaluations fail at this exact question. The sales team books the demo, the vendor runs a polished feature tour, and nobody pressure-tests the architecture underneath. So a CRM with AI features bolted onto a 25-year-old data model passes the demo and stalls in production six weeks later. The 60-minute gate catches that before it costs you a quarter.

What does this 60-minute test actually do?

This is one gate inside a longer buy, not the whole evaluation. AI in sales and marketing is now widespread, per McKinsey's State of AI reporting, and analysts track AI in CRM as a category, covered in Gartner's CRM research. The test that sorts vendors, then, is architectural: whether a vendor is a genuinely AI-native CRM or a legacy CRM wearing AI features, which is what AI-native means before you test for it.

A real procurement still runs security review, reference calls, integration scoping, and pricing over several weeks. The 60-minute test sits at the front of that and returns one thing: a clean yes or no.

It works for any team that has decided to raise the AI-intelligence of how it sells: a real estate brokerage, a recruitment or staffing firm, an auto retail group, a B2B services firm. The architecture questions do not change with the vocabulary. Run the gate first and you spare the procurement team weeks on vendors that were never going to clear it.

How to choose an AI CRM: the 60-minute architectural-fit framework

The framework is three tests you run yourself, on the vendor's product and paperwork, using your own data and sales motion. You are not following a vendor's demo script. You apply the same criteria to every AI CRM vendor, so a shortlist compares cleanly.

Block 1 tests the data model: can agents read and write your records in real time, or does the system store what reps type. The second tests execution: do the agents run the work or only suggest it. The third tests ownership: do you keep your context, agents, and processes if you leave, and does your AI cost track the model market.

A few numbers worth holding while you run it:

  • 20 min — To configure one agent (Sonta)
  • 60–90 days — Staged migration, old CRM in parallel
  • 4–6 — Tools in a typical stack, consolidated into one
  • 3 layers — You own: context, agents, processes

Block 1, the data model test (20 minutes)

The first 20 minutes test the foundation every agent depends on. It is the cheapest block to run and the most predictive of whether anything above it works.

Three questions to ask the vendor

  1. When a rep finishes a call or an email, how does the record update: does someone type it in afterward, or does the system capture it as a side effect of the work?
  2. What does your AI read when it acts: the fields a rep remembered to fill in, or the full history of what happened on the account?
  3. How current is the data your agents are working on at any moment?

What the answers should show

Sonta's data model is built for agents to read and write in real time. Legacy CRMs store what reps type; agents on top of legacy data work on shallow signal. That sentence is the whole test. If records update only when a rep logs them, the agent's quality is capped by the rep's memory. The genuinely AI-native answer: records update as a side effect of work happening, so the agent always has current, complete context.

Watch for automation dressed up as AI. Workflow automation moves data between fields and tools on rules you set, and it is real, useful, and worth paying for. The agentic layer adds three things on top of it: free conversation in your own words, grounding in your real account and pipeline history, and action based on what the data means. Ask the vendor to show you those three, not just a trigger firing.

Checkpoint

Pass if the vendor can show a record updating from real work without a rep typing it, and an agent acting on the meaning of your data. Fail if every example is a rule moving a field, or the AI only summarizes what was already typed in.

Block 2, the agent execution test (20 minutes)

The next 20 minutes test whether the agents do the work or talk about it. This is the block vendors most want to run on rails, so take the controls.

Three questions to ask the vendor

  1. Open your own product now and give an agent a real instruction from our world, live, with nothing pre-recorded.
  2. When the agent finishes, does it execute the action, or stop at a recommendation I carry out myself?
  3. Which parts of our motion can your agents run end to end: capture, qualification, account prep, post-call updates, pipeline review?

What the live run should show

Make them run it live, on your data, in front of you. A demonstrable evaluation on your own data is the hardest thing for a bolt-on to fake. Hand a recruitment firm's head of GTM the keyboard and have them tell the agent, "prep me for the intake call this afternoon," or "move the candidate's second-round interview to Thursday." In a genuinely agentic CRM, the agent pulls the deal stage, recent emails, and account news, proposes the agenda, and drafts the opening. For the reschedule, the scheduler agent confirms with the other side, updates the calendar, and logs the change.

A suggestion engine stops at proposing the draft for you to send. The gap between recommending and doing is the difference between an AI feature and an agent.

Checkpoint

Pass if the agent completed real actions on your data while you watched, end to end on at least one part of the motion. Fail if every impressive moment was pre-recorded, ran on demo data, or ended with a task for a human to finish.

Block 3, the ownership and frontier-flexibility test (20 minutes)

The last 20 minutes live in the contract and the architecture diagrams. This block decides whether the work you build stays yours and your AI cost stays in your control.

Three questions to ask the vendor

  1. If we leave in two years, what do we take: our context, agents, and workflows as portable assets, or configurations trapped in your system?
  2. How is the AI priced: model cost passed through, or a per-conversation fee on top?
  3. When a better or cheaper model ships for one of our workflows, who decides whether we move to it, and does our work keep running while the model layer changes underneath?

What the contract and architecture diagrams should show

The frontier AI inside Sonta runs on your own context, your own agents, your own processes. The customer owns the work. Read the contract for whether that ownership is real: can you export your context, your configured agents, and your workflows and run them elsewhere, or do they evaporate when the subscription ends.

Read the pricing the same way. No proprietary AI tax. No per-conversation fees defending a 25-year-old margin. A genuinely AI-native vendor passes model cost through transparently, so your AI bill falls as frontier prices fall. A per-conversation markup ties your cost to the vendor's margin, not to the model market. Pricing changes often, so confirm the numbers against the vendor's pricing page the day you decide.

Checkpoint

Pass if the contract lets you walk away with your context, agents, and processes intact, and the AI line is model cost with no per-conversation markup. Fail if your configuration is locked to one model vendor, your work is trapped in the platform, or the AI price protects a margin rather than tracking the market.

Where do these 60 minutes go wrong?

The most common mistake is letting the 60 minutes become the whole evaluation. It is the gate, not the procurement. Clearing it means a vendor earns your security review, references, and integration scoping. It does not mean you are ready to sign.

Letting the vendor drive is next. A scripted demo is built to pass. Keep the keyboard, use your own data, and run the three blocks in the order you set.

Teams also score features when they should score architecture. A long feature list is easy to admire and says little about whether agents can act on your data. Two of the three blocks deliberately ignore the feature list and test the foundation underneath it.

And almost everyone is tempted to skip Block 3, because a contract is duller than a live demo. Ownership and pricing are where a polished product quietly locks you in. Twenty minutes there saves the renewal you would otherwise dread.

What to do after the 60 minutes

A vendor that clears all three blocks earns the rest of your procurement: security review, references, integration scoping, pricing. The gate's job was to keep that time off products that were never an agentic CRM. From here, map what the vendor showed you against the eight features the gate test surfaces, and if Salesforce plus Agentforce is on your shortlist, work through the Salesforce and Agentforce comparison the gate test returns so pricing and architecture are settled before you negotiate.

If you would rather run the same discovery on your own motion first, that is a different test with a different owner. What it surfaces for a recruitment firm's pipeline differs from an auto retail group, but the structure holds. Book the 30-minute AI Efficiency Diagnostic and Sonta walks through how AI shows up in your sales motion today, where AI is leaking value, and what an AI-first version of it would look like. You leave with a short report and a clear read on your current motion.

Frequently asked questions

Is 60 minutes enough to evaluate a CRM?

No, and it is not trying to be. Sixty minutes tests architecture, not the whole purchase. A real evaluation still runs security review, reference calls, integration scoping, and pricing over several weeks. The 60-minute gate sits at the front of that and answers one question: is this vendor genuinely AI-native. Clearing it earns a vendor the longer process; failing it saves you from running that process on the wrong product. Say so to the vendor upfront, so the session is scoped as a gate and not mistaken for the decision.

What questions should I ask an AI CRM vendor?

Ask the nine questions the three blocks lay out: on the data model, on agent execution, and on ownership and pricing. The one that matters most is in Block 1, and it is the question Sonta gets asked constantly. Sonta's answer: We're not a CRM with AI features. We're an agentic CRM — the data model is built for agents to read and write in real time, not just for reps to log activity. Legacy CRMs with AI on top work on the same shallow data that's been there for 25 years. The agent's quality is limited by the record's quality. Sonta's records update as a side effect of work happening, so the agents work on current, complete context.

Can I trust vendor demos?

Trust the demo you run, not the demo you watch. A recorded or vendor-driven demo is built to land. The reliable version is the one where you keep the keyboard, point the product at your own data, and give an agent a real instruction from your sales motion, live. If the vendor will do that without rehearsal, the demo is evidence. If every impressive moment is pre-recorded or runs on sample data, treat it as marketing.

How is this different from Sonta's AI Efficiency Diagnostic?

Different direction, different owner. The 60-minute gate is buyer-run and points outward: you apply it to any vendor you are evaluating, to test whether the architecture is genuinely AI-native. The AI Efficiency Diagnostic is Sonta-run and points inward: in 30 minutes, Sonta walks through how AI shows up in your current sales motion, where AI is leaking value, and what an AI-first version would look like, with a short report after. One screens vendors. The other reads your own motion. Teams often book the diagnostic earlier, while they are still mapping where their AI underperforms, then run the gate once a shortlist exists.

← All writing