Data Science Lifecycle

The data science lifecycle follows CRISP-DM, the standard process for data science projects. myBytes works in its six phases and walks through an example of what happens in each phase.

From Your Raw Data
to Business Impact

A model is only as robust as the process that produces it. That is why we skip none of the six phases.

CRISP-DM

CRISP-DM: Definition, Origin and How It Differs

The Cross-Industry Standard Process for Data Mining describes in six phases how a business question becomes a model in operation.

CRISP-DM stands for Cross-Industry Standard Process for Data Mining. The process model breaks a data science project down into six phases: Business Understanding, Data Understanding, Data Preparation, Modeling, Evaluation and Deployment. Each phase has a guiding question, defined tasks and a deliverable that the next phase needs. Since the 2000s it has been regarded as the de facto standard for data mining and data science projects.

CRISP-DM was developed from 1996 by Daimler-Benz, NCR and the tool vendor ISL, which later became part of SPSS, with the insurer OHRA as the fourth partner; the CRISP-DM 1.0 guide was published in 2000. The model is deliberately industry-neutral and tool-independent. That is exactly why it fits projects nobody had in mind in 1996: a returns model in fashion retail, a failure forecast in manufacturing and a turnover model in logistics all pass through the same six phases.

The newer process models build on CRISP-DM rather than replacing it. ASUM-DM from IBM adds operations and project management, TDSP from Microsoft adds roles, tasks and a project structure but has not been maintained since 2023, and DASC-PM from Nordakademie extends it with data provisioning and scientific grounding. Anyone who knows one of these models will recognise CRISP-DM in it.

CRISP-DM itself does not define roles; in practice the phases still fall to recognisable owners. Phases 1 and 2 need the business side: the business question, the success criteria and the knowledge of which data exist in-house and what a field in the merchandise system actually means come from the people who work with the data. Phases 3 and 6 are data engineering: pipelines, versioning, operations. Phases 4 and 5 are the work of the data scientists. No phase belongs to one role alone, and a project that staffs only the middle phases has neither a robust question nor an operation.

CRISP-DM is iterative, not linear. The arrows also run backwards. If the data assessment in phase 2 shows that the business goal from phase 1 cannot be answered with the available data, the project goes back to phase 1. If a model misses the success criteria in evaluation, it goes back to modelling or to data preparation. The outer loop closes after deployment: new data from operations raise new business questions.

Data science lifecycle

The Complete Lifecycle

Each phase builds on the previous one: iterative, transparent and aligned with your business goals.

Data science phases

The Six Phases of the Data Science Lifecycle in Detail

Click on a phase for the full description.

Phase 01

Business Understanding

Defining business goals and translating them into analytical questions.

Guiding question: Which decision should be better at the end, and how will you measure that? Deliverable: a target picture with success criteria and an assessment of whether the data can carry the question.

Goal DefinitionSuccess CriteriaFeasibility
View details →
Phase 02

Data Acquisition & Exploration (Data Understanding)

Identifying and accessing relevant data sources with quality assessment.

Guiding question: Which data do you already have in-house, and is it enough for the question from phase 1? Deliverable: a data inventory with a quality rating per source.

Data SourcesQuality AnalysisExploration
View details →
Phase 03

Data Processing & Engineering (Data Preparation)

Cleaning, transforming and deriving meaningful features.

Guiding question: How do raw data become clean, reproducible datasets? Deliverable: versioned data pipelines with quality checks. This phase takes up the largest share of project work.

Data CleaningFeaturesAutomation
View details →
Phase 04

Exploratory Analysis & Modeling

Systematic model development and optimisation with traceable experiments.

Guiding question: Which model answers the question better than a simple reference? Deliverable: documented experiments with a baseline and the best candidates.

AnalysisModel DevelopmentOptimisation
View details →
Phase 05

Evaluation & Validation

Rigorous assessment of model performance against your business goals.

Guiding question: Does the model meet the success criteria from phase 1 under realistic conditions as well? Deliverable: a validation report with metrics, driving factors and a release recommendation.

Performance ReviewExplainabilityValidation
View details →
Phase 06

Deployment & Monitoring

Production deployment with monitoring and continuous improvement.

Guiding question: How does the model get into operation, and who notices when it degrades? Deliverable: a production system with monitoring and an operations handover to your team.

OperationsAutomationMonitoring
View details →
CRISP-DM example

CRISP-DM Example: Fashion Returns, Worked Through in Six Phases

A use case from our fashion series, phase by phase. All figures are a model calculation for a fictitious company, not client results.

A mid-sized fashion company with 1,800 articles and 18 markets sells jeans online. The return rate for jeans is 38 percent, and 34 percent of those returns are down to "does not fit".

Phase 1 · Business Understanding

The business question is not "we want AI" but: What share of fit-related returns can be avoided if the customer is recommended the right size at checkout? The success criterion is the return rate per category, measured before and after the launch. Phase 1 also sets the limit: new customers without a purchase history can initially only be served with category averages by such a model.

Phase 2 · Data Acquisition & Exploration

The data are in-house: purchase history, return reasons, the ordered and the kept size per customer and cut, plus the size charts per article. Phase 2 checks whether the size information in the returns data is complete. If it is missing, there is no model, only a statistic.

Phase 3 · Data Processing & Engineering

From the raw data, each customer gets a profile of kept and returned sizes per cut, and each article a calibrated size run. The model calculation assumes that a correct size recommendation would have been enough for 68 percent of fit-related returns. Whether that assumption holds is decided here, on the quality of the size fields.

Phase 4 · Exploratory Analysis & Modeling

The model is collaborative filtering with body-shape matching: customers with similar purchase and return behaviour are recommended the size that similar customers kept for this cut. The benchmark is the size chart without a recommendation, as a simple reference model before any more elaborate approach. If the model does not beat that reference, the effort is not worth it.

Phase 5 · Evaluation & Validation

In the model calculation the model picks the right size in 82 percent of cases. Fit-related returns fall by 18 percent, and the jeans return rate falls from 38 to 31 percent. A published vendor case reaches up to 20 percent fewer total returns; the model calculation deliberately stays below that. The evaluation also checks whether the effect appears in all categories or only for jeans.

Phase 6 · Deployment & Monitoring

The model runs as a widget on the product page: "Based on your previous purchases we recommend size 31/32." A second output goes to the product team: articles whose size chart deviates systematically. Monitoring keeps watching the return rate per category. The model calculation puts 61,824 avoided returns at 26 euros each, which is 1.61 million euros a year.

PhaseGuiding question in the exampleDeliverable in the example
1 Business UnderstandingWhat share of fit-related returns is avoidable?Success criterion: return rate per category
2 Data UnderstandingAre size and return reason recorded per order?Data inventory from purchase history, returns, size charts
3 Data PreparationHow do customer and cut become comparable?Size profile per customer, calibrated size run per article
4 ModelingWhich size do similar customers keep?Collaborative filtering with body-shape matching
5 EvaluationDoes the model pick the size more often than the chart?82 percent hit rate, return rate 38 to 31 percent (model calculation)
6 DeploymentHow does the recommendation reach the customer?Widget on the product page, report to the product team, monitoring

The next step in the same series shows the second iteration of the loop: a text analysis of 210,000 free-text comments breaks "does not fit" down into five causes and, together with the size recommendation, lowers the overall return rate of the model calculation across all categories and 18 markets from 28 to around 23 percent.

Reducing return rates: the use case behind the example Analysing return reasons: what lies behind "does not fit"
Data science project

Where Data Science Projects Fail - and in Which Phase

The published figures on failure point to the first two phases, not to modelling.

The MIT NANDA programme reported in August 2025 that 95 percent of the organisations studied see no measurable effect on results despite their spending on generative AI. A year earlier, Gartner had predicted that at least 30 percent of GenAI projects would be abandoned after the proof of concept by the end of 2025, because of poor data quality, missing risk controls, rising costs or unclear business value. In our view, unclear business value and missing risk controls are questions of phase 1, data quality a question of phase 2; the costs usually only show in operations. Our research article traces the seven methodological causes, from claim inflation in the project proposal to comparison against a weak baseline.

Why 95% of GenAI pilots fail

The second finding concerns the data. In our conversations with mid-sized data teams, AI pilots rarely fail on the model and mostly on a data quality nobody checked beforehand. For that we have published an audit that checks a dataset against seven dimensions in ten minutes, with a traffic-light rating per dimension. The audit belongs in phase 2, before effort flows into phase 3.

Seven questions decide whether your AI project fails

The third finding is the simplest: without a written goal there is nothing a model could pass or fail against. The countermeasure costs little in phase 1: a written, pre-registered success endpoint that says which metric is measured on which population over which period against which comparison. Without that endpoint nothing can be validated in phase 5, and the project ends up being judged on impressions instead of numbers.

When a project ends at myBytes, it usually ends after phase 1 or phase 2, not only after modelling. That is intentional. A no after the target picture or after the data assessment is cheap; a no after modelling comes only once the most laborious phase, data preparation, is already done. If your data cannot carry the question, we say so and propose what needs to happen first.
Data science methodology

Our Data Science Approach

Three principles run through every phase of our projects.

Iterative & Agile

Short feedback loops with continuous stakeholder involvement.

Reproducible & Transparent

Documented decisions and traceable results.

Business-First

We measure success by real value delivered, not by model complexity.

CRISP-DM phases

Frequently Asked Questions about CRISP-DM

What does CRISP-DM mean and what does the abbreviation stand for?

CRISP-DM stands for Cross-Industry Standard Process for Data Mining. The process model was developed from 1996 by Daimler-Benz, NCR and ISL (later SPSS) and published as a guide in 2000. It describes six phases from the business question to operations and is industry-neutral and tool-independent by design.

What are the six phases of CRISP-DM?

Business Understanding, Data Understanding, Data Preparation, Modeling, Evaluation and Deployment. At myBytes they are called "Business Understanding", "Data Acquisition and Exploration", "Data Processing and Engineering", "Exploratory Analysis and Modeling", "Evaluation and Validation" and "Deployment and Monitoring". The phases run iteratively: stepping back into an earlier phase is intended, not an exception.

Is CRISP-DM still current, or are there newer process models?

Newer models such as ASUM-DM, TDSP and DASC-PM extend CRISP-DM but do not replace it. The 2000 guide treats operations and monitoring only as a task within deployment. That is exactly where we extend phase 6 with monitoring, automated release and operations handover. The six phases themselves endure.

How long does a data science project take under CRISP-DM?

There is no blanket duration, and we deliberately do not name one. The runtime depends on whether the data from phase 2 can carry the question and how much preparation phase 3 demands. Two things can be said reliably: phase 1 starts with a conversation, and a first data check against seven dimensions takes ten minutes.

Which phase takes the most time?

Data preparation in phase 3. It takes up the largest share of project work because missing values, duplicates and inconsistencies from several sources come together here. That is why we automate it with reproducible pipelines and quality checks. What that looks like in detail is on the page for phase 3: data preparation.

What is the difference between CRISP-DM and the data science lifecycle?

CRISP-DM is a concrete, documented process model with six named phases. "Data science lifecycle" is the umbrella term for any cyclical representation of a data science project; many versions are variants of CRISP-DM with a stronger emphasis on operations. Our lifecycle is CRISP-DM under our own phase names, with monitoring as part of the sixth phase.

Is there a CRISP-DM example from practice?

Yes. On this page we work through fashion returns across all six phases, from the business question to the running model, with the guiding question and the result for each phase. The use case behind it shows data, method and business case.

Can CRISP-DM also be used for AI projects with language models?

Yes, and the first two phases become more important there, not less. The published reasons for abandoning GenAI projects, unclear business value and poor data quality, are questions of phase 1 and phase 2. With language models, modelling usually means selection and integration rather than training your own; evaluation and monitoring remain mandatory as before.

The first step towards data-driven results

Every project starts with phase 1, understanding your business. That first step is a conversation, not a contract.

Every project is unique. Tell us about your challenge.

Start your free initial call