Godwit AI Labs Talk to us

HomeServices › Data & AI

Data and AI, without the demo

Most of the AI ideas we are asked to build do not need AI, and most of the ones that do are blocked by data that does not exist yet. We will tell you which of yours is which, then build the two or three that pay for themselves.

ADOPTION · WEEK SIX illustrative 68% of eligible decisions routed through it week 1 week 8 12% held for human review 3% overridden by the operator
What adoption looks like when somebody is counting · illustrative, not a client’s numbers
Fee
Readiness review, fixed fee + GST · 3 weeks
Runs for
3 weeks
You provide
Access to the data, and the people who use it
Output
Two or three use cases, costed
01

The unglamorous part comes first

Talk to us about this

Every AI project that failed quietly in the last two years failed for the same reason: the data was not there, was not trusted, or lived in four systems that disagreed with each other. A model cannot answer a question about numbers nobody can reconcile.

  • A warehouse or lakehouse that is the one place a number is allowed to come from.
  • Pipelines out of the systems that matter (ERP, CRM, billing, the spreadsheets that turn out to be load-bearing) on a schedule, with failures that alert somebody.
  • Definitions written down, so that "active customer" means one thing across finance and sales rather than two.
  • Reporting people open: dashboards built where your team already works, in Power BI, Looker Studio or Zoho Analytics. This is often the whole return on the project before any AI is discussed.

If your real problem is that three departments each have a different revenue number, we will say so, fix that, and stop.

Mostof the AI ideas we are asked about do not need AI
Firstthe data has to exist somewhere a system can reach
Two or threeuse cases is what a mid-market company can land in a year
02

Which use cases are worth building

Talk to us about this

We score candidates on four things and publish the workings: is the data there, is the decision valuable, is the output checkable, and what happens when it is wrong. Most ideas fail on the third or the fourth.

  • Usually worth it. Document extraction over a real backlog, search across internal knowledge nobody can find, classification and routing of inbound volume, forecasting where a decent forecast already exists to beat.
  • Usually not. A chatbot on a website with eleven pages, anything whose output nobody downstream is allowed to trust, and anything where the honest baseline is a database query somebody could write this afternoon.

A demo is easy and proves almost nothing. The hard part is what happens on the four hundredth document, with a bad scan, when the person who checked the first ten has moved on.

03

Agents, and where they fit

Talk to us about this

An agent is a different proposition from a model that answers a question: it takes actions. It reads the mailbox, updates the record, raises the ticket, calls the next system. That difference is the entire risk. A wrong answer is embarrassing; a wrong action has already happened.

  • Where they earn their keep. Multi-step back-office work with a checkable result and a small blast radius: matching an invoice against a purchase order and a goods receipt, chasing the document that has not arrived, triaging inbound requests to the right queue with the reason attached.
  • Where they do not. Anything irreversible without a person in the loop. Anything acting on a system with no audit trail, because you will not be able to reconstruct what happened. And anything a scheduled job would do more cheaply and far more predictably, which is a surprising share of what gets asked for.

When we do build one: permissions scoped to the task rather than inherited from whoever configured it, every action logged and reversible, and a stop condition, because an agent that cannot finish should stop and say so rather than improvise. The failure we design against is not one bad answer. It is twenty confident bad actions before anybody looks.

04

Finding the use cases in the first place

Talk to us about this

Most AI shortlists get written in a meeting room by the people furthest from the work. We would rather go and watch the work: where the queue builds up, what gets re-keyed from one screen into another, which decision waits on one person being available, and what people keep in a spreadsheet because the system will not let them do their job.

On a shop floor that means the line, not the ERP screen: the checks a supervisor makes by eye, the rejection that gets logged four hours after it happened, the handover at shift change where half the context goes home with the person leaving. That is where the measurable cases are, and none of them turn up in a strategy deck.

You get the shortlist with the workings, including the rejections: two or three worth building, and for everything dropped, the reason it was dropped.

05

Training and adoption, which is where most of this fails

Talk to us about this

The model is rarely why an AI project does not land. It does not land because the people it was built for carried on doing it the old way, and nobody was counting.

  • Train the people who do the work, not the managers who sponsored it, on their shift, on their screens, in the language they work in.
  • On a shop floor, respect the floor. Gloves, noise, no desk, a shared terminal, a supervisor covering three lines. Anything needing a laptop and a quiet minute will not get used, and that is a design failure rather than a discipline problem.
  • Show people what it gets wrong. Trust comes from knowing the failure mode, not from being quoted an accuracy figure. A tool nobody trusts gets quietly worked around, and you find out months later.
  • Measure adoption, not deployment. How many decisions went through it last week. If that number is falling we treat it as a defect in the thing we built, not as a training problem belonging to you.

We stay through that stretch rather than handing over at go-live, because the first month decides whether any of it was worth building.

06

Then we build the ones that survive

Talk to us about this

Small, in your accounts, and instrumented so you can see whether it is working.

  • Built into a workflow somebody already uses, rather than as a separate tool to remember.
  • Evaluated against a set of real cases, with accuracy stated as a number rather than as an impression, and a threshold below which it does not ship.
  • A human check on the decisions that carry consequences, and a log of what the system did.
  • Running cost measured from the first week, because inference is a bill that grows with success.
07

Guardrails, before anything ships

Talk to us about this

This is where data work and the DPDP Rules meet, and where a lot of enthusiasm should stop for a moment.

  • Which personal data is going into a model, and whether the purpose you collected it for covers this one.
  • Where the processing physically happens, and what your contracts say about that.
  • What is retained: prompts and outputs are records too, and they are discoverable.
  • A written position on what staff may paste into public tools, which most companies need far more urgently than they need a model.

We build this alongside the security work, because the answer usually depends on your data map either way.

08

Then somebody has to run it

Talk to us about this

A model is not finished when it ships, for the same reason a migration is not finished when the last workload copies across. Four things go wrong afterwards, and not one of them announces itself.

  • The data underneath moves. Somebody renames a column, a source system is upgraded, and a pipeline that was correct on Friday is quietly wrong on Monday. The dashboard still renders a number. That is the problem.
  • Accuracy drifts. The cases it sees in month six are not the ones it was evaluated on. Without a standing evaluation set and somebody re-running it, nobody finds out until a person acts on a bad answer.
  • The bill grows with success. Inference is charged by use, so the better it works the more it costs. That needs a budget alert and a monthly read, not an annual surprise.
  • Nobody opens it. A report that has stopped being used is a running cost with nothing coming back, and the fix is usually a conversation rather than a rebuild.

So we run it: the pipelines, the evaluations, the guardrails and the cost, on the same managed service as the rest of the estate. If you would rather hold it yourselves, the handover is the runbook, the evaluation set and the dashboards, in your accounts. We will say so plainly when that is the better answer.

DELIVERABLES

What you walk away with

Yours to keep, and to hand to anyone else, including a provider that isn't us.

Talk to us about this
An honest shortlist

The two or three use cases worth money, and the reasons the others were dropped.

A data readiness read

What exists, what is trustworthy, and what would have to be built before anything works.

A cost and a payback

Build cost, running cost, and the number the case has to beat.

A guardrail note

What your own data policy and DPDP obligations allow before anything ships.

QUESTIONS

Asked before you ask

The ones that come up on nearly every first call.

Talk to us about this
We do not have a data warehouse. Is that a blocker?

For reporting, no. We can often get useful answers from the source systems while the platform is built. For anything predictive, usually yes, and it is better to hear that in week one than after a proof of concept that could never have gone to production.

Can you just build the chatbot the board asked for?

We can, and we will first show you what it will and will not be able to answer given the content you have. If the honest answer is that a decent search box and three rewritten pages would serve people better, that is what we will recommend.

Do you build your own models?

No. We use hosted models and open-weight ones, chosen per use case, and we tell you which so that you are not locked to our preference. Training something from scratch is almost never the right answer at mid-market scale.

Does our data get used to train someone else's model?

Not under the enterprise terms we would put you on, and we will show you the specific clause rather than reassure you. This is the question your customers will eventually ask you, so it is worth having the paperwork straight.

How is this different from the analytics our BI vendor sells?

Often it is not, and where a BI tool already fits we will say so rather than rebuild it. Our work tends to start where a dashboard stops being enough: when the decision needs to be made on volume, at speed, or over unstructured content.

Not sure this is the one you need?

Tell us the symptom, not the service, and we'll say which of these applies, or that none of them do. A reply within one working day, from someone technical.

Talk to us