Buyer's guide
How to choose an AI development company in Canada
A practical checklist for buyers: what a production AI system needs, what to ask a vendor, and which answers should end the conversation. Use it on us too.
What AI development should mean for a buyer
Plenty of companies now list AI development as a service, and the label covers very different work. At one end is a demo that answers three scripted questions. In the middle is a model API behind a chat box. At the other end is a system that real users rely on every day, with your real data behind it. You are paying for the last one, so say so in the first conversation.
A production AI system has parts a demo skips: your actual data, a way to measure whether the output is right, limits on what the model may do, a plan for when it is wrong, and someone who keeps it healthy after launch. If a vendor can talk about models for an hour but cannot explain those five, you are buying a demo.
Decide what you need, too. Many useful features are an existing model API plus retrieval over your own documents, careful evaluation and a good interface. Training a custom model is the exception. A vendor who proposes it first should explain why an API model is not enough.
An evaluation checklist for any AI development company
Use these eight checks with every vendor on your shortlist, including us. For each one, ask for a concrete answer you could verify.
- Shipped production AI. Ask for AI products that real people use today, and open them yourself. A live product tells you more than a slide. Then ask what went wrong after launch and what changed. Teams that have run AI in production have a list.
- Evaluation of model output. Ask how you will know the output is good enough before launch. A solid answer describes an evaluation set built from real examples in your domain, a result for every change to a prompt, model or retrieval step, and a record of results over time. Checking by hand is not an evaluation method.
- How they handle your data. Ask where your data is stored and processed, which model providers see it, how long anyone keeps it, whether it can train any model, and who on their side can access it. Get the answers in writing. For personal information, PIPEDA is Canada's federal private-sector privacy law and provincial rules may also apply, so raise your privacy obligations early and take advice on what applies to you.
- Cost of ownership. An AI feature has a running cost: every request to a model provider is billed by usage, and that grows with your users. Ask how cost scales with volume, what keeps it in check (caching, smaller models for easy tasks, limits per user), and who owns the provider accounts and sees the bills.
- Who actually writes the code. Ask who will build your system, how long they have done production work, and whether the person who scopes the project is the person who builds it. Then ask what happens if someone leaves midway. Continuity matters on AI work, because the reasons behind prompt and retrieval choices live in people's heads unless written down.
- Code and IP ownership. The code, prompts, retrieval pipelines, evaluation sets, any fine-tuned model and the infrastructure should belong to you from the first commit. Ask where the repository lives and whose accounts everything runs in. If the answer is theirs, leaving later will be expensive.
- Handover and documentation. Ask what you will hold at the end: a readable repository, tests, a README that gets a new engineer running, and notes on the non-obvious decisions. Could another team pick the project up on Monday without calling the vendor?
- Post-launch support. Models change, providers retire versions and real inputs drift away from your test cases. Ask who watches quality after launch, how a regression gets caught, and how ongoing maintenance is arranged.
Red flags that should slow you down
Any one of these is a reason to ask more questions. Two or three together is usually a reason to walk away.
- Accuracy promised before they have seen your data. Nobody can quote a reliable accuracy figure for your problem until they have tested on your examples. A confident number up front is a sales line.
- No evaluation plan. If quality depends on the model being state of the art, or on someone trying it a few times, you will find the failures after your users do.
- A portfolio of demos. Screenshots are not products. Ask for a live link and the person who runs it.
- Your work living in their accounts. Their repository, their cloud, their model keys. That is lock-in, whether or not anyone intended it.
- Vague answers about data. “It is secure” and “it is encrypted” are not answers. Where, who, how long and for what purpose are.
- No plan for wrong answers. Models are sometimes confidently wrong. A good design says what happens then: a fallback, a limit, or a handoff to a person.
Questions to ask before you sign
Put these to every vendor in writing.
- Which AI products have you shipped that people use today, and can I open them?
- How will we measure whether the output is good enough to launch, and who builds that test set?
- Where will my data be stored and processed, and which third parties will see it?
- Can my data be used to train any model, yours or a provider's?
- How does the running cost of model usage grow as we grow, and what keeps it under control?
- Who will write the code, and will that person stay on the project from scoping to launch?
- Who owns the repository, prompts, evaluation sets and cloud accounts, and from when?
- What will I receive at handover, and could another team take over from it?
- What happens when the model gets an answer wrong in production?
- Who looks after quality after launch, and how is that work arranged?
A vendor with real experience answers in specifics. One without it answers in adjectives.
How Devino answers these
Our answers, limited to what we already publish on this site.
- Shipped AI. We build and run our own AI products. Shorty summarizes YouTube videos and Spotify podcasts, SuperBooks categorizes bank transactions and matches them to receipts, and GetItDone gives AI agents the same task context a person has.
- Evaluation. Before tuning a prompt, we turn real examples from your domain into an evaluation set and measure every prompt, model or retrieval change against it. SuperBooks runs a golden dataset of matching cases in its test suite, so a drop in quality shows up before it ships.
- Your data. Your data stays in your infrastructure wherever possible. We avoid sending sensitive data to third-party models unnecessarily, and we document what leaves your systems and why.
- Who builds it. Senior engineers only. Every engineer brings 5+ years of production experience and trains for six months inside Devino before a first client project. The engineer who scopes your work is the one who builds it.
- Ownership and visibility. You own the code from day one, including prompts, retrieval pipelines and evaluation sets, in a repository with tests and docs. You get a deployed link every week to try on your own cases.
- Cost and life after launch. We build AI systems to survive real users, real data and real costs, not just the demo. The cloud and model-provider accounts are registered to you and billed to you, so you see every usage bill yourself. Once a product is live, maintenance is available as hourly support.
Already have an AI demo or a stalled build? Our software consulting starts with an audit and an honest read, and the same engineers can ship the fix. To start a new project, describe it through the intake form and an engineer replies within one business day.
Frequently asked questions
Have an AI project in mind?
Describe it through the intake form. An engineer replies within one business day with a scoped plan and an honest read on fit, not a sales call.
Get it built