Modelling, data mining and pipelines: predict, overlap and inspect
Move beyond method definitions. Calculate what a data pattern suggests, predict a batch's completion time under stated assumptions, and choose a representation that helps someone make a decision.
Content owner: Michael Print · Written for A-Level learners · Checked against official specifications
The idea to start with
Data mining finds potentially useful patterns in collected data; a useful decision also needs context, comparison and validation. Performance modelling predicts time or resource use under explicit assumptions. Pipelining overlaps different stages for different items while preserving each item's stage dependencies.
Visualisation presents data or a model so relationships can be inspected. It can reveal waiting, changing demand or diminishing returns, but it does not establish that the underlying data is representative or the model is correct.
OCR H446 · 2.2.2(f): data mining, performance modelling, pipelining and visualisation
Before you start
Useful foundations
Percentages, averages and basic arithmetic
Independent work, dependencies and concurrent thinking
Valid predictions depend on assumptions
By the end, you should be able to
Interpret an observed pattern against its baseline and limitations.
Build and check a simple performance model.
Schedule dependent stages for multiple items and identify the bottleneck.
Choose a labelled visualisation that exposes the relationship of interest.
Mine a pattern, then ask whether it helps
An invented café examines 120 transaction baskets: 30 contain coffee, 60 a sandwich and 24 both. It asks whether a sandwich suggestion beside coffee merits a trial.
Data mining seeks useful relationships, such as co-occurring purchases. Compare the conditional rate with its baseline; a popular item naturally has many co-occurrences.
The pattern proves neither that coffee causes sandwich purchases nor that suggestions increase sales. Time of day, duplicates or a busy exhibition weekend could distort the result.
Clean the data, define the population and check a later separate set. A trial should measure additional suitable purchases while considering irrelevant suggestions and waste.
These are fictional aggregate baskets. Actual data mining still requires appropriate privacy and lawful use; a technical correlation does not settle those decisions.
Original basket counts and interpretation
Measure
Calculation
Meaning
Sandwich baseline
60 / 120 = 50%
Rate across all baskets
Sandwich given coffee
24 / 30 = 80%
Rate within coffee baskets
Relative rate
0.80 / 0.50 = 1.6
Observed rate exceeds the baseline
Causal effect
Not established by these counts
Needs a suitable separate investigation
Turn a workflow into a performance model
Each image passes validation, transformation and export. With constant stage times and no transfer, startup or failure overheads, sequential time is 2+4+3=9 seconds per image: 36 seconds for four.
Give each stage a worker and buffer intermediate images. A single image still follows every dependency, but different workers handle different images at once.
The first image needs nine seconds. Transform is the bottleneck, so later images complete four seconds apart. For n≥1, predicted batch time is 9 + (n - 1) × 4; zero images require a separate zero-duration result.
The equation is a performance model; stage overlap is pipelining. One overloaded disk or one worker doing every stage can defeat the assumptions. Compare measurements with predictions before adding workers.
Steady-state capacity is fifteen images per minute. Persistent arrivals at twenty per minute grow the backlog. Buffers absorb bursts, but cannot fix a sustained rate mismatch; finite buffers eventually fill.
An image’s route through the pipeline
1
Validate: 2 seconds
Worker 1 checks the input before it may enter transform.
2
Transform: 4 seconds
Worker 2 performs the slowest stage. It limits the steady-state completion interval.
3
Export: 3 seconds
Worker 3 writes the transformed output.
Stages for one image remain ordered. Different images can occupy different workers simultaneously.
Use visualisation to answer a specific question
A timeline puts seconds on one axis and workers on separate rows. Label each image’s interval. The schedule table is an accessible equivalent: B waits after validation while transform finishes A.
Compare like with like: keep workload, stage rules and units consistent. Label axes, samples and measured/predicted values. A truncated time axis can exaggerate an improvement.
Keep data and explanations alongside colour. Readers should not need to distinguish colours to understand the conclusion. A visible pattern still needs checks for data quality and causal interpretation.
Choose a visual for the question
Where does work wait?→
Use a worker timeline with labelled image intervals.
Does throughput match the model?→
Plot elapsed time against completed images, with measured and predicted series.
Which configuration finishes sooner?→
Compare batch durations using bars under the same workload.
Does image size affect processing time?→
Use a scatter plot of size against actual transform time.
Worked example
Schedule four images without violating their dependencies
A validates from 0 to 2, transforms from 2 to 6, then exports from 6 to 9. Its completion latency is nine seconds.
B validates while A transforms, from 2 to 4. B must wait until transform becomes free at 6; it then transforms 6–10 and exports 10–13.
C validates 4–6, transforms 10–14 and exports 14–17. D validates 6–8, transforms 14–18 and exports 18–21.
The batch finishes at 21 seconds, agreeing with 9 + 3 × 4. Completion times 9, 13, 17, 21 have a four-second interval after the first image. The first image's latency is unchanged even though batch duration falls from 36 to 21.
The transform worker is busiest. Improving export alone from three seconds to one would lower startup latency but leave the four-second steady-state bottleneck until another stage becomes slower.
Original pipeline schedule; intervals in seconds
Image
Validate
Transform
Export
Completion
A
0–2
2–6
6–9
9
B
2–4
6–10
10–13
13
C
4–6
10–14
14–17
17
D
6–8
14–18
18–21
21
Worked example
Predict a targeted improvement and state its limits
Add a second transform worker, assuming transforms are independent and the workers have no shared-resource contention. Each image still takes four seconds in transform, but two can overlap.
The transform stage's aggregate capacity becomes one image per two seconds. Validation also has a two-second capacity, while export still needs three seconds per image. Export becomes the bottleneck.
For this constant-time ordered workload, the first image completes at 9 seconds and later images complete every 3 seconds. Four images finish at 18 seconds; six finish at 24. Draw or calculate the schedule before generalising this model to variable stage times.
Measure the revised system. If exports compete with transforms for storage access, the predicted improvement may not occur. A performance model helps choose an experiment but does not replace its evidence.
Original A-Level practice
6 original questions total 25 marks. Attempt each before opening the independently written indicative marking guidance.
Question 1
4 marks
In 200 invented baskets, fruit appears in 80, yoghurt in 100 and both in 60. Calculate the yoghurt baseline, yoghurt rate within fruit baskets and the ratio of those rates. Give one interpretation limit.
Show solution and marking guidance+
Indicative answer
Baseline is 100 / 200 = 50% (1). Within fruit baskets it is 60 / 80 = 75% (1). The ratio is 0.75 / 0.50 = 1.5 (1). This observed association does not establish causation or guarantee future behaviour; accept another specific data-quality limit (1).
Question 2
4 marks
Use the one-worker-per-stage image model for six images. Calculate sequential and pipelined duration and state two assumptions supporting the prediction.
Show solution and marking guidance+
Indicative answer
Sequential duration is 6 × 9 = 54 seconds (1). Pipeline duration is 9 + 5 × 4 = 29 seconds (1). Award one mark each for two relevant assumptions, such as separate stage workers, fixed durations, sufficient buffers, ignored transfers or no shared-resource contention (2).
Question 3
4 marks
Explain why B cannot begin transformation at time 4 in the displayed schedule. Give B's completion time and distinguish this waiting from the first image's latency.
Show solution and marking guidance+
Indicative answer
Transform is occupied by A until time 6 (1), so B waits despite having completed validation (1). B completes at time 13 (1). A's nine-second first-item latency is its path through the stages; pipeline waiting for later items is a separate scheduling effect (1).
Question 4
4 marks
The original pipeline receives twenty images per minute continuously. Calculate its sustainable processing rate, explain the buffer consequence, and name the stage a first capacity experiment should target.
Show solution and marking guidance+
Indicative answer
Its four-second bottleneck gives 60 / 4 = 15 images per minute (1). Arrival exceeds completion by five images per minute under the model (1), so backlog grows and finite buffers eventually fill (1). Target transform, the slowest stage (1).
Question 5
4 marks
Choose a visualisation for locating pipeline waiting and another for checking whether image size predicts transform duration. State two presentation choices that support a sound interpretation.
Show solution and marking guidance+
Indicative answer
A labelled worker/stage timeline exposes occupied and waiting intervals (1). A scatter plot of image size against measured transform time explores that relationship (1). Award one mark each for two useful choices such as units, clear scales, measured/predicted labels, workload/sample information or an accessible accompanying data table (2).
Question 6
5 marks
A café manager says the observed coffee/sandwich pattern proves that displaying a suggestion will increase sales. Give two reasons the claim is unsupported and design a way to assess the proposed decision.
Show solution and marking guidance+
Indicative answer
Co-occurrence does not establish that suggestions cause purchases (1). Confounding, unrepresentative sampling or duplicate/bias errors can explain the observed pattern (1). Define the relevant outcome, such as additional suitable sandwich purchases (1). Compare an appropriate controlled trial or another justified evaluation using new data (1), monitoring limitations or unwanted effects such as irrelevant suggestions/waste (1).
Specification and references
This guide addresses OCR H446 2.2.2(f): data mining, performance modelling, pipelining and visualisation. Check your examination year and the complete specification for the assessment scope.
These are independently written explanations and practice questions. CompSciTutoring.co.uk is not affiliated with or endorsed by an examination board. The marking guidance is indicative; always check the syllabus for your examination year.