Explain a processor by following the information it needs, then identify what makes that work faster. A performance answer needs a mechanism and a workload, not just the word faster.
Content owner: Michael Print · Written for A-Level learners · Checked against official specifications
The idea to start with
The control unit coordinates instruction execution; the ALU performs arithmetic and logic; registers hold the small amount of information needed immediately. Address, data and control buses connect these operations to memory and other components.
Clock speed sets the rate of clock cycles, cache reduces some memory-access delays, and additional cores can execute independent work simultaneously. None alone guarantees that every program finishes sooner.
OCR H446 · 1.1.1(a–e)
Before you start
Useful foundations
Binary values and memory addresses
The existing fetch–decode–execute guide
By the end, you should be able to
Trace a load and an addition through named registers
Evaluate clock, cache and core changes in context
Trace an ideal pipeline and explain its limitations
Compare Von Neumann, Harvard and contemporary designs
Give each component a precise job
Real processors may have many general-purpose registers rather than one accumulator. The diagram uses the simplified accumulator model for the worked trace.
An instruction contains an operation and, where needed, an operand. For a direct-address load, the operand identifies a memory location whose contents are required. The control unit decodes the operation, issues control signals and coordinates transfers.
The ALU can add, subtract, compare and perform operations such as AND. A comparison may affect status flags used by a later conditional branch.
Which register holds which information?
PC → next instruction address→
The program counter identifies the instruction to fetch next.
MAR → memory address→
The memory address register identifies the location to read or write.
MDR → transferred data→
The memory data register holds an instruction or value moving to or from memory.
CIR → current instruction→
The current instruction register holds the instruction being decoded or executed.
ACC → operand or result→
The accumulator holds arithmetic operands and results in this simplified machine.
An address names a location; its contents are the value stored there.
Bus roles in a simplified processor
Bus
What travels
Example
Address
Location to access
Address 20 sent from MAR to memory
Data
Instructions or operand values
Value 7 read into MDR
Control
Signals coordinating operations
Read signal; write signal; interrupt request
Connect the fetch cycle to an assembly instruction
In a typical school model, PC is copied to MAR; the control unit requests a memory read; the instruction travels into MDR and then CIR; PC advances to the next instruction.
Exact physical timing varies, so state the model instead of claiming every processor copies registers in an identical order.
For LDA 20, fetching obtains the instruction first. Executing it requires a second memory access: MAR becomes 20, MDR receives the contents of location 20, then ACC receives that value.
For ADD 21, location 21 supplies an operand to the ALU, which adds it to ACC. A branch replaces the next-instruction address in PC. Do not put an operand's value into MAR when an address is required.
Clock frequency: cycles per second
A higher clock frequency offers more cycles per second. If an otherwise identical CPU executes the same number of cycles without additional delays, it finishes sooner. Different instruction sets, instructions per cycle, thermal limits and memory waits make clock-only comparisons unreliable between different processors.
Raising frequency can increase power consumption and heat.
Cache: reduce repeated memory waits
Cache holds copies of recently or frequently used instructions/data close to the processor. A cache hit can avoid a slower main-memory access; a miss still requires lower-level memory.
Larger cache may improve a program that repeatedly reuses a working set, but gives little benefit to some streaming workloads. Cache is distinct from registers and from the disk that stores files.
Cores: identify independent work
Several cores can run different processes or parallel parts of one process. If the next calculation depends on the previous result, additional cores cannot simply divide that sequence evenly. Communication, synchronisation and shared-memory contention reduce ideal speed-up. Identify the independent work before recommending more cores.
Pipelining overlaps stages
In a three-stage model, one instruction is fetched while an earlier one is decoded and an earlier one is executed. The pipeline increases throughput after filling; it does not reduce the three-stage latency of one instruction.
A data dependency may require a stall, and a branch can invalidate instructions fetched along the wrong path. Filling, draining and hazards explain why the ideal rate is not always achieved.
Compare memory architectures
Von Neumann architecture uses a shared memory/address space for instructions and data, with a shared transfer path in the simple model. An instruction fetch and data transfer compete for that path.
Harvard architecture separates instruction and data memories and their buses, allowing suitable instruction/data accesses concurrently. Separate stores can also need different widths or capacities.
Contemporary processors often combine ideas: separate instruction/data caches near a core, unified main memory further away, multiple execution units and several cores. This is often called a modified Harvard arrangement.
Describe the actual separation in the scenario; do not infer that every modern CPU is purely one historical architecture.
Worked example
Trace two instructions
Assume one address per instruction, PC initially 0, ACC initially 0, memory[0] contains LDA 20 and memory[1] contains ADD 21. Data locations 20 and 21 contain 7 and 5.
The fetch of LDA uses instruction address 0; execution uses data address 20. ACC becomes 7. Fetching ADD uses address 1; execution uses address 21 and ACC becomes 12. PC now points to 2.
State after executing each instruction; CIR contains the fetched instruction
Instruction
PC
MAR after operand read
MDR after operand read
ACC
LDA 20
1
20
7
7
ADD 21
2
21
5
12
Worked example
Four instructions through an ideal pipeline
Assume four independent instructions, one cycle per stage and no hazards. Without overlap, four × three = 12 cycles. With overlap, the total is three fill cycles plus three further completions = six cycles.
The ideal steady state completes one instruction per cycle, but the first instruction still takes three cycles. A real branch or dependency would require revising the trace.
F = fetch, D = decode, E = execute, dash = inactive
Cycle
I1
I2
I3
I4
1
F
—
—
—
2
D
F
—
—
3
E
D
F
—
4
—
E
D
F
5
—
—
E
D
6
—
—
—
E
Original A-Level practice
4 original questions total 15 marks. Attempt each before opening the independently written indicative marking guidance.
Question 1
3 marks
A direct-address instruction loads memory location 42 containing 19. Distinguish MAR, MDR and ACC during execution.
Show solution and marking guidance+
Indicative answer
1 mark: MAR holds the address 42.
1 mark: MDR receives value 19 from memory.
1 mark: ACC receives 19 for subsequent processing.
Question 2
4 marks
Six independent instructions use three one-cycle stages. Calculate cycles with and without ideal pipelining and explain one reason actual time may be longer.
Show solution and marking guidance+
Indicative answer
1 mark: without pipelining, 6 × 3 = 18 cycles.
1 mark: ideal pipeline needs 3 + 5 = 8 cycles.
1 mark: identify a hazard, such as an operand dependency or mispredicted branch.
1 mark: explain that waiting or flushing adds cycles.
Question 3
4 marks
A program repeatedly processes the same small lookup table but performs each step sequentially. Evaluate larger cache and more cores.
Show solution and marking guidance+
Indicative answer
1 mark: reuse of the table makes cache hits plausible.
1 mark: more hits reduce slower main-memory reads, if the table fits and was not already cached.
1 mark: dependent sequential steps cannot simply execute simultaneously on several cores.
1 mark: conclude cache is the more relevant improvement here, while other independent processes could still benefit from extra cores.
Question 4
4 marks
Compare Von Neumann and Harvard designs and describe a contemporary combination.
Show solution and marking guidance+
Indicative answer
1 mark: Von Neumann shares instruction/data memory and its path in the simple model.
1 mark: Harvard separates instruction/data memories and paths.
1 mark: this separation can permit simultaneous instruction and data access.
1 mark: a contemporary combination is separate instruction/data caches with unified main memory.
Specification and references
This guide addresses OCR H446 1.1.1(a–e). Check your examination year and the complete specification for the assessment scope.
These are independently written explanations and practice questions. CompSciTutoring.co.uk is not affiliated with or endorsed by an examination board. The marking guidance is indicative; always check the syllabus for your examination year.