(Data science)

Data science that ends with a decision someone acts on

Most of data science is the work around the model: framing the question, building features that hold up, and proving the thing beats the rule of thumb it replaces. Fitting the model is rarely the hard part.

Data mesh abstract
Data science

(Before anyone fits a model)

A question worth modelling has a decision attached to it. Which claim to route to an assessor, which order is likely to come back, how much stock to hold in the northern depot next month. Without that decision, a model is a curiosity with a good score, and the team holding it has no way to tell whether it helped.

Sometimes the honest answer is a rule and a threshold rather than a model. That is cheaper to run and far easier to explain, and you will hear it before the modelling budget is spent.

(What we deliver)

What data science consulting covers here

Engagements usually begin at the top of this list. Each step is worth stopping after if the evidence says so.

  1. Turning a business ambition into a prediction with a target, a unit of analysis and a time horizon. Includes the awkward questions: is the label available at prediction time, and does anyone act differently once they have the answer?

  2. The variables that carry the signal, built from your transactional data with a careful eye on leakage. Rolling windows, aggregates and joins are written as code in a repository, not derived by hand in a notebook nobody can rerun.

  3. Gradient boosting for tabular problems, time-series methods for demand and load, deep learning where text or images genuinely require it. Model choice follows the data and the latency budget rather than fashion.

  4. A/B tests, holdout groups and uplift modelling to answer whether the change caused the outcome. Sample size and stopping rules are agreed before the test starts, so results cannot be read early and declared a success.

  5. Performance broken down by segment, cost of a false positive weighed against a false negative, and a catalogue of the cases the model gets wrong. Calibration matters as much as accuracy when a number drives a decision.

  6. A packaged model with a defined input contract, a reproducible training pipeline, a registry entry and monitoring for drift. Your engineers get something they can deploy without reading a research notebook.

(Usual stack)

(How we work)

How a data science project runs

Five phases with a decision point after each. The early ones are deliberately cheap, because that is where most projects should stop.

  1. 01

    Frame

    A week with the people who make the decision today. The output is a written problem statement: what is predicted, for whom, when, and what changes as a result.

  2. 02

    Data audit

    How much history exists, how clean it is, and whether the label is trustworthy. Plenty of ideas die here, and dying here costs a fortnight instead of a quarter.

  3. 03

    Baseline

    The current process, or a simple rule, measured properly on the evaluation set. Every model afterwards has to beat it by enough to justify the cost of running it.

  4. 04

    Iterate

    Features, models and tuning in short cycles, with results logged in MLflow so you can see what was tried. Progress is shown as a moving metric, not as a promise.

  5. 05

    Evaluate and hand over

    Error analysis by segment, a calibration check, sign-off against the measure agreed in framing, then packaging, documentation and a live pilot with a fallback to the old process.

(Why Team of Keys)

Why our data science work holds up after handover

Models rot quietly. The habits below are about making that visible rather than pretending it does not happen.

  1. 01

    The baseline is a real contender

    A rule, an average or the existing human process is measured on the same evaluation set. If it wins, the project stops and you keep the analysis and the saving.

  2. 02

    Leakage is hunted, not assumed away

    Features are checked against what was genuinely known at prediction time. A model that looks excellent in testing and fails in production is almost always a leakage story.

  3. 03

    Reproducible from raw data

    Training runs from a scripted pipeline with pinned dependencies and versioned data. Anyone on your side can rebuild the model six months later without the original author.

  4. 04

    Drift is monitored

    Input distributions and output rates are watched after launch, with an alert when they move. Retraining has a trigger and an owner rather than being remembered after complaints.

  5. 05

    Explanations people accept

    Feature importance and per-case explanations for any decision that affects a customer or a member of staff, so the team using the model can defend an individual outcome to the person on the receiving end.

(Related)

More in ai & data

All ai & data services

(FAQ)

Questions, answered

Framing and the data audit are a fixed, small price and take two to three weeks between them. Modelling is quoted per phase once we know what the data supports. A single prediction problem with clean history is much cheaper than one needing a new pipeline first, and you will know which you have before committing.

It depends on the problem more than the row count. A monthly forecast needs several years of history to see seasonality; a churn model might work with tens of thousands of labelled accounts. The audit answers this directly, and where the data is thin we will say so rather than fitting something unstable.

Only where text is the input. For tabular prediction, gradient boosting remains more accurate, cheaper to run and far easier to explain. Language models earn their place in extraction, classification of free text and summarisation, and in those cases the evaluation set matters even more.

You are told, with the numbers. That result usually arrives in the baseline phase, before serious money is spent, and it is still useful: you leave with a measured view of the existing process and a clear account of what data would need to change for a model to help.

Either your engineers or us. Handover includes the training pipeline, the registry entry, monitoring dashboards and a runbook for retraining. Where you want us to keep it running, that sits in an agreed monthly capacity covering drift reviews, retraining and dependency updates.

(Global presence)

Nine countries, one studio behind them.

Every project is designed, built and shipped from one studio.
Turn the globe, or pick a country to see what we deliver there.

(Next step)

Bring us the decision, not the dataset

Tell us what your team decides today by instinct, and what data sits behind it. You get a framing note and a phased price, usually within two working days.

START

Or write to info@teamofkeys.com · Noida, India