Cambridge 9618 gives AI both an AS applications/impact context and an A Level technical treatment. Explain what the system learns, what evidence trains it and how its output is evaluated.
Content owner: Michael Print · Written for A-Level learners · Checked against official specifications
The idea to start with
Supervised learning uses labelled examples to learn a mapping; unsupervised learning seeks patterns without supplied target labels; reinforcement learning learns action choices using rewards from interaction. Deep learning uses neural networks with multiple learned layers and can support different learning setups.
Training changes model parameters to improve a defined objective. Performance on training data alone is insufficient: unseen evaluation data, bias, costs and human consequences matter. An AI output is not automatically accurate because it was computed.
Cambridge International 9618 · 9618 (2027–2029): 7.1 AI applications/impact; 18.1 graphs, neural networks, learning categories, backpropagation and regression
Before you start
Useful foundations
Graphs and weighted paths
Functions, averages and simple algebra
By the end, you should be able to
Match learning methods to data and a task
Explain neural-network layers, weights and error feedback
Perform a small regression prediction/update
Use graph-search ideas without claiming graph implementation is required
Evaluate social, economic and environmental effects
Keep the examination scope explicit
In 2027–2029 version 2, AS section 7.1 covers AI applications and social, economic and environmental impacts. The checked 2026 version 2 retains these requirements.
A Level section 18.1 adds A*/Dijkstra graph search, neural networks, machine/deep/reinforcement learning, supervised/unsupervised categories, backpropagation and regression. This technical scope extends beyond OCR’s ethical discussion of AI.
Cambridge explicitly exempts writing graph setup/access/search algorithms for 18.1. Understand graphs and perform the named searches. Shared pathfinding reasoning is useful; another board’s graph-code assessment requirements must not be imported.
Choose the learning method from its evidence
Learning methods use different training evidence. Match that evidence to the task before selecting a model, and decide how to measure success on unseen examples.
Supervised training reduces differences between predictions and supplied targets. Labels can be expensive or biased, and a training sample can omit rare cases. An unsupervised cluster indicates similarity, rather than a verified fault diagnosis.
A reinforcement agent seeks a policy improving expected reward. Reward design, exploration and costly or unsafe trials limit practical use; a useful reward must reflect the behaviour actually wanted.
What information teaches the system?
Supervised
Input features plus known targets: classify labelled fault photographs or regress expected energy consumption.
Unsupervised
Inputs without supplied target labels: cluster similar sensor patterns, then interpret what each group means.
Reinforcement
An agent acts in an environment, changes state and receives rewards: learn a robot’s route choices through trials.
Deep learning describes the network
Deep learning uses multiple learned layers that can build useful representations, such as image features. It does not exclude supervised learning: a deep network can learn from labelled images.
Larger networks can require substantial data and computation. A smaller, interpretable model may suit a simple numerical relationship better. A named AI application does not establish that AI is the best approach.
A neuron weights, adds and activates
A neuron weights input values, adds a bias and applies an activation function. Connected layers pass outputs forward; the learned parameters determine the network’s mapping.
With inputs 2 and 1, weights 0.4 and −0.2, and bias 0.1, the weighted sum is 0.4×2−0.2×1+0.1 = 0.7. ReLU is defined here as max(0,z), giving output 0.7.
Training uses error to update parameters
Backpropagation uses the chain rule to calculate each parameter’s contribution to loss through the layers. An optimisation method then updates parameters; backpropagation itself does not copy the desired answer into every weight.
The learning rate controls update size. Too large may destabilise learning; too small may slow progress. One update does not guarantee a perfect prediction.
A supervised training cycle
1
Predict
Pass training inputs forward through the network using the current weights and biases.
2
Measure loss
Compare the prediction with its supplied target using the chosen loss function.
3
Backpropagate
Calculate how changes to each parameter would affect the loss.
4
Update
Use an optimisation method and learning rate to change the parameters, then repeat with training examples.
Improvement on training examples still needs independent evaluation on unseen data.
Regression predicts quantities
Linear regression predicts y using y-hat = w×x+b. Fitting chooses w and b to reduce training loss, rather than treating the first guessed line as truth.
Mean squared error averages squared prediction errors, strongly penalising large errors. Regression is a prediction method; a fitted relationship is not evidence that one measured variable causes the other.
Evaluate generalisation and who bears errors
Use separate held-out validation/test data. Repeatedly choosing model changes from test results leaks information into development; a final untouched test provides stronger evidence.
Overfitting captures training-specific patterns but performs poorly on new inputs. Check representative sampling, class imbalance and false-positive/false-negative consequences. High overall accuracy can hide poor performance for a small important group.
An automated fault detector can reduce manual work and catch faults early; false alarms interrupt production and missed faults create risk. Workers may gain tasks or face displacement; data collection can expose personal information.
Training and deployment consume energy. Consider benefits, error costs and the roles of oversight, appeal, staff training and proportionate data use. A reliable fixed threshold may meet a simple task with lower complexity.
Graph search compares route costs
Graphs represent states or locations as vertices, possible moves as edges and costs as weights. Dijkstra expands minimum cost-so-far choices for non-negative weights.
A* prioritises estimated total f=g+h: travelled cost plus a heuristic. A suitable non-overestimating heuristic supports shortest-path guarantees under the algorithm’s stated rules. The worked trace below shows repeated cost updates without graph implementation code.
Worked example
One regression training step
Original training points: (1,4) and (2,7). Start with w=2 and b=1. Predictions are 3 and 5, errors are −1 and −2, and mean squared error is (1+4)/2=2.5.
For this explicitly defined squared-error model, gradient for w is (2/2)×[(−1)×1+(−2)×2]=−5; gradient for b is (2/2)×[−1−2]=−3. At learning rate 0.1, subtracting the gradients gives w=2.5 and b=1.3.
New predictions are 3.8 and 6.3; squared-error average is (0.04+0.49)/2=0.265. This improves the training objective but supplies no independent test-set evidence. Gradient arithmetic is supporting depth; the syllabus's backpropagation understanding does not require implementing a training library.
Prediction/error before and after the update
x
Target y
Old prediction
Old error
New prediction
New error
1
4
3
-1
3.8
-0.2
2
7
5
-2
6.3
-0.7
Runnable Python 3: a defined one-step regression modelpython
points = [(1.0, 4.0), (2.0, 7.0)]
w, b = 2.0, 1.0
errors = [w * x + b - y for x, y in points]
loss = sum(error ** 2 for error in errors) / len(points)
grad_w = 2 * sum(error * x for error, (x, y) in zip(errors, points)) / len(points)
grad_b = 2 * sum(errors) / len(points)
w -= 0.1 * grad_w
b -= 0.1 * grad_b
new_loss = sum((w * x + b - y) ** 2 for x, y in points) / len(points)
print(f"before: {loss:.3f}")
print(f"w={w:.1f}, b={b:.1f}")
print(f"after: {new_loss:.3f}")
Worked example
Search an original graph without implementing it
Directed edges are S→A (2), S→B (5), A→B (1), A→G (6), B→G (2). Costs are non-negative; start at S and finish at G.
Dijkstra settles S, A, B, G. Follow the distance updates in the table, then use predecessors to reconstruct S→A→B→G with cost 5.
For A*, use h(S)=5, h(A)=3, h(B)=2, h(G)=0. Each estimate is no larger than the true remaining distance.
After S, A has g=2,f=5 and B has g=5,f=7. After A, B improves to g=3,f=5 and G receives g=8,f=8.
After B, G improves to g=5,f=5. A* reaches G with the same shortest route. Its priority is g+h, not h alone.
If edges change, check the heuristic again: an old estimate can overestimate the new shortest distance. This trace does not need graph implementation code.
Dijkstra tentative distances after expanding each vertex
Expanded
A
B
G
S
2
5
∞
A
2
3
8
B
2
3
5
G
2
3
5
Worked example
Choose a learning method
Known historical energy readings with measured targets support supervised regression. Labelled fault photographs support supervised classification, potentially using a deep network if data/resources justify it.
Unlabelled machine telemetry can support unsupervised clustering, but someone must interpret whether a cluster represents a relevant fault. A robot adjusting actions from trial rewards is reinforcement learning.
For any option, define the success measure, test unseen representative cases and consider the cost of errors. A reliable fixed threshold may be preferable if it meets the task with lower complexity.
Original A-Level practice
7 original questions total 24 marks. Attempt each before opening the independently written indicative marking guidance.
Question 1
3 marks
Match three tasks to supervised, unsupervised or reinforcement learning: labelled disease images, grouping unlabelled activity logs, and an agent choosing actions from rewards.
Show solution and marking guidance+
Indicative answer
1 mark: labelled images use supervised learning/classification.
Explain a neuron's weighted input, bias and activation, then calculate the output of this neuron.
Neuron inputs and parameters
Bias is 0.1. Add the bias to the sum of the weighted inputs to obtain z, then use ReLU activation, defined as max(0,z).
Inputs to the neuron
Input
Value
Weight
First
1
0.4
Second
3
-0.2
Show solution and marking guidance+
Indicative answer
1 mark: weights scale each input's contribution.
1 mark: the bias offsets the combined weighted sum.
1 mark: activation transforms that sum; here ReLU=max(0,z).
1 mark: z=0.4×1−0.2×3+0.1=−0.1, so ReLU output is 0.
Question 3
3 marks
Explain the roles of loss, backpropagation and learning rate during training.
Show solution and marking guidance+
Indicative answer
1 mark: loss measures mismatch according to the chosen objective.
1 mark: backpropagation calculates parameter gradients through the network.
1 mark: learning rate sets the size of the parameter update; it does not guarantee a correct answer.
Question 4
3 marks
A regression model predicts y=3x+1. For points (1,5) and (2,6), calculate predictions and mean squared error.
Show solution and marking guidance+
Indicative answer
1 mark: predictions are 4 and 7.
1 mark: errors are −1 and 1, giving squared errors 1 and 1.
1 mark: mean squared error is (1+1)/2=1.
Question 5
3 marks
A model achieves perfect accuracy on its training data. Give two reasons this is insufficient and one improvement to evaluation.
Show solution and marking guidance+
Indicative answer
1 mark: it may have overfit training-specific patterns.
1 mark: the training set may be unrepresentative/imbalanced or omit important cases.
1 mark: evaluate separately on held-out representative data, including performance for important subgroups/error costs.
Question 6
4 marks
Evaluate an AI recruitment filter with one benefit, two risks and one practical safeguard.
Show solution and marking guidance+
Indicative answer
1 mark: consistent fast initial processing can reduce staff workload.
1 mark: biased/unrepresentative training labels can disadvantage applicants.
1 mark: opaque errors, unnecessary personal-data collection or energy/computation costs can harm stakeholders.
1 mark: appropriate oversight, subgroup evaluation, an appeal route or data minimisation addresses a named risk; explain the connection.
Question 7
4 marks
The graph below uses a changed S→B cost of 1. Give the shortest route/cost, calculate B's initial A* priority using h(B)=2, and assess the old h(S)=5.
Graph and heuristics for this question
All edges are directed as listed. Start at S; the goal is G. The old heuristics are h(S)=5, h(A)=3, h(B)=2 and h(G)=0. For A*, use f = g + h, where g is the cost travelled from S.
Edges after changing S→B
From
To
Cost
S
A
2
S
B
1
A
B
1
A
G
6
B
G
2
Show solution and marking guidance+
Indicative answer
1 mark: S→B→G now has cost 1+2=3 and is shortest.
1 mark: B has g=1 and f=g+h=1+2=3.
1 mark: h(S)=5 overestimates the new remaining distance 3.
1 mark: update/check the heuristic before claiming admissibility/shortest-path guarantees, or use h=0 to recover Dijkstra's priority rule.
Specification and references
This guide addresses Cambridge International 9618 9618 (2027–2029): 7.1 AI applications/impact; 18.1 graphs, neural networks, learning categories, backpropagation and regression. Check your examination year and the complete specification for the assessment scope.
These are independently written explanations and practice questions. CompSciTutoring.co.uk is not affiliated with or endorsed by an examination board. The marking guidance is indicative; always check the syllabus for your examination year.