Dulranga's Notes
Semester 3Computer Architecture

Computational Limitations

1. Power wall problem

Why we don't see computers with 10GHz clock speeds? Only around 3.2GHz in normal life. That's not because of just expensive to make them, its because of a fundamental limit with computation.

In a CPU, Higherf∝More heat\text{Higher} f\propto \text{More heat} At speeds as even 5-6 GHz, CPUs dissipate massive amount of heat that needs to be removed before melting the chip.

CMOS digital circuits:

Pdynamic=α⋅C⋅V2⋅fP_{\text{dynamic}} = \alpha \cdot C \cdot V^2 \cdot f

Where:

  • PdynamicP_{\text{dynamic}} = Dynamic power dissipated as heat
  • α\alpha = Activity factor (percentage of transistors switching per cycle)
  • CC = Capacitance of the circuits
  • VV = Operating supply voltage
  • ff = Operating clock frequency

Total power is the sum of dynamic power and static leakage power:

Ptotal=Pdynamic+PstaticP_{\text{total}} = P_{\text{dynamic}} + P_{\text{static}}

To keep dynamic power same when increasing ff, we need to lower voltage. But there is a limit we can lower, after that the voltages are not enough for transistors to switch effectively. Voltage was lowered by decreasing the size of the transistor. But again after some time, the size cannot be reduced more due to the silicon starts to leak more current which increases PstaticP_{\text{static}}

At some point, the power usage is not worthy enough for increasing clock speed. This is known as the Power wall Problem

2. Dark Silicon

This is also a direct result of power wall problem. see the α\alpha factor. which is how many transistors are switching directly contribute to the power dissipation. To reduce power usage, we keep most of the transistors switched off completely (ie. In the dark)

At modern process nodes (3nm, 5nm), up to 50% to 90%+ of a chip's silicon may need to remain "dark" depending on the workload.

More Ways to reduce Power dissipation

Modern CPUs use a multi core architecture with optimal clock speeds for each core to reduce power usage. Further optimizing this, Mac M1 series processors offer two type of cores,

  1. Performance cores Designed for maximum throughput with wider execution pipelines and higher clock speeds (~4.0+ GHz). However, pushing P-cores to peak frequency requires significantly higher supply voltage (VV), causing power draw to scale exponentially right up against the Power Wall.

  2. Efficient cores Built with narrower pipelines, lower operating frequencies (~2.0–2.5 GHz), and much lower voltages (VV). Because voltage is squared in the power equation, E-cores consume roughly 1/10th the power of P-cores for light tasks.

Fixed-Function Accelerators

Apple further addresses Dark Silicon by allocating massive chip real estate to specialized accelerators instead of general-purpose CPU cores:

  • Neural Engine (NPU): Specialized for AI/ML inference.
  • Media Engine: Hardware units for H.264, HEVC, and ProRes video decoding/encoding.
  • GPU & Matrix Cores (AMX): Specialized for graphics and linear algebra math.

When exporting a 4K video, the general-purpose CPU cores remain mostly idle ("dark"), while the hyper-efficient Media Engine turns on to encode the video using a tiny fraction of the power a CPU would require.

On this page