115 年 國立中山大學資訊工程學系碩士班甲組《計算機結構》
第 Q1.1 題7 分
- Short Conceptual Questions (20 points)
Q1.1 (7 points)
For several decades, processor performance improvements were largely achieved by increasing clock frequency. However, this approach eventually stopped being the primary driver of performance gains. Modern processors now rely heavily on techniques such as instruction-level parallelism, pipelining, speculative execution, and hierarchical memory systems.
Please explain why increasing clock frequency alone has become an ineffective and unsustainable method for improving processor performance. In your explanation:
- (4 points) Identify two fundamental physical or architectural constraints that limit the benefits of higher clock frequencies.
- (3 points) Explain how each constraint directly motivates the architectural shift toward parallelism, prediction, and memory hierarchy.
(Your answer should be based on core computer architecture principles rather than manufacturing or marketing considerations.)
登入後即可作答並保存紀錄。
核心觀念
處理器執行時間可表示為:
其中 為指令數、 為每指令平均週期數、 為時脈週期、 為時脈頻率。
提高 只能縮短單一時脈週期;若 與 不變,效能才會近似按比例提升。實際處理器的 會受到資料相依、分支錯誤、快取未命中與主記憶體等待影響,因此頻率提高不代表所有瓶頸都同步改善。
本題應從兩個核心限制切入:
- 功耗與散熱上限。
- 訊號與記憶體延遲造成的記憶體牆。
解題方法一:功耗與散熱上限
動態功耗可近似表示為:
其中 為切換活動率、 為負載電容、 為電壓、 為頻率。當頻率提高時,動態功耗至少隨 增加;為了讓電路在更短週期內完成運算,通常還需要較高電壓,導致功耗以更快速度上升。同時,漏電功耗與溫度也會增加,最終受到晶片可承受的功率與散熱能力限制。
因此,提高頻率會產生以下問題:
- 功耗與溫度急遽增加。
- 時脈網路與管線暫存器本身也消耗大量功率。
- 為了維持安全溫度,處理器必須降頻,頻率提升的效益因而消失。
- 更深的管線雖能縮短每級邏輯延遲,卻會增加管線暫存器、控制複雜度與錯誤分支的損失。
這項限制直接促成三種架構方向:
- 平行性:採用超純量(superscalar)、多重發射、亂序執行(out-of-order execution)或多核心,在適度頻率下同時處理多個獨立工作,提高每週期完成的指令數,即提升 。
- 預測與推測執行:管線越深,分支錯誤時必須清除的管線級數越多。分支預測(branch prediction)先猜測方向與目標,推測執行(speculative execution)沿預測路徑工作,使管線保持充滿,減少控制冒險造成的空轉。
- 記憶體階層:暫存器與快取可利用資料區域性,減少高延遲主記憶體存取,也降低處理器長時間等待及不必要的資料搬移能量。
解題方法二:訊號延遲與記憶體牆
處理器的時脈週期不能任意縮短,因為必須容納最長組合邏輯路徑及暫存器的時序需求:
晶片內長距離導線、時脈分配與不同運算單元之間的訊號傳遞,都具有實際延遲。當時脈越快,這些固定的傳遞時間相當於越多個時脈週期,因此不能藉由單純提高頻率消除。
更嚴重的是記憶體牆(memory wall):處理器頻率提升後,主記憶體的實際存取時間並未按相同比例縮短。若主記憶體延遲為固定的奈秒數,換算成處理器週期後:
第 Q1.2 題7 分
Q1.2 (7 points)
A program executes on a processor with the following characteristics:
• Total instruction count: instructions
• Instruction mix:
* 50% arithmetic instructions, CPI = 1
* 30% memory instructions, CPI = 5
* 20% control instructions, CPI = 2
• Processor clock frequency: 2 GHz
Please answer the following questions:
- (3 points) Compute the total CPU execution time of the program.
- (4 points) The processor designer improves the memory system so that the CPI of memory instructions is reduced from 5 to 2. Compute the new execution time and the overall speedup.
登入後即可作答並保存紀錄。
核心觀念
本題考查處理器效能的基本計算:
總 CPU 週期數為各類指令所需週期數的總和:
也可先計算加權平均 CPI:
再利用:
解題方法
題目給定總指令數、各類指令比例與 CPI,因此先分別計算三類指令的數量,再計算總週期數,最後除以時脈頻率得到執行時間。
1. 原始處理器的 CPU 執行時間
總指令數為:
各類指令數量如下:
因此總 CPU 週期數為:
處理器時脈頻率為:
所以 CPU 執行時間為:
2. 改善記憶體系統後的執行時間與加速比
改善後,只有 memory instructions 的 CPI 從 降為 ,其他條件皆不變。
新的總週期數為:
第 Q1.3 題6 分
Q1.3 (6 points)
Many architectural optimizations are designed to "make the common case fast." Please explain what this principle means in practice and why violating it often leads to worse overall performance, even if rare cases become faster.
登入後即可作答並保存紀錄。
核心觀念
本題考查「讓常見情況變快」(make the common case fast)的設計原則,以及平均效能與 Amdahl’s Law。
在處理器中,各種事件發生的頻率不同,例如:
- Cache hit 通常比 Cache miss 常見。
- TLB hit 通常比 TLB miss 常見。
- 正常的簡單指令通常比複雜例外處理常見。
- 分支預測正確通常比預測錯誤常見。
設第 種情況的發生比例為 ,處理時間為 ,則平均處理時間為:
因此,架構設計不應只看某一種情況的最佳延遲,而要看所有情況依發生比例加權後的總體效能。
「讓常見情況變快」在實務上表示:將硬體資源、控制邏輯與關鍵電路優先配置給最常發生的執行路徑,使大多數指令或記憶體存取能以最短、最直接的路徑完成;罕見情況則保留正確但較慢的處理路徑。
解題方法
設常見情況發生比例為 ,執行時間為 ;罕見情況發生比例為 ,執行時間為 。
原本的平均執行時間為:
假設某項設計使罕見情況加速 倍,但由於增加硬體複雜度,使常見情況多花費 ,則新的平均執行時間為:
要使整體效能改善,必須滿足:
整理可得:
左側是所有常見情況共同承受的額外成本;右側是罕見情況加速所節省的時間。當 很小時,右側通常很有限,因此只要常見路徑增加少量負擔,就可能抵銷罕見路徑的全部收益。
例如,某處理器中:
- 的情況屬於常見路徑,每次耗時 個週期。
- 的情況屬於罕見路徑,每次耗時 個週期。
原本平均耗時為:
若為了讓罕見路徑從 個週期降至 個週期,卻使常見路徑增加 的成本,則:
雖然罕見情況變快了,但平均耗時由 增加至 ,整體效能反而下降約 。
第 Q2.1 題10 分
- Performance Analysis (20 points)
Q2.1 (10 points)
A program executes on a processor with the following characteristics:
• 35% of instructions are arithmetic instructions (CPI = 1)
• 45% are memory access instructions (CPI = 3)
• 20% are control instructions (CPI = 2)
• Clock cycle time = 0.4 ns
The program executes 1 billion instructions.
Please calculate:
- (5 points) Average CPI
- (5 points) Total CPU execution time
登入後即可作答並保存紀錄。
核心觀念
本題考查指令比例、加權平均 CPI,以及 CPU 執行時間的計算。
平均 CPI 不能直接將各類指令的 CPI 取普通平均,而必須按照各類指令所占比例加權:
CPU 執行時間公式為:
解題方法與計算
1. 計算平均 CPI
三類指令對平均 CPI 的貢獻分別為:
- 算術指令:
- 記憶體存取指令:
- 控制指令:
因此:
平均而言,每一條指令需要 個 clock cycles。
2. 計算總 CPU 執行時間
指令總數為:
所需的總 clock cycles 為:
第 Q2.2 題10 分
Q2.2 (10 points)
An architect accelerates a floating-point unit so that floating-point instructions run five times faster. Originally, floating-point instructions account for 30% of total execution time. Please use the fundamental performance laws to compute the maximum possible speedup of the entire program.
登入後即可作答並保存紀錄。
核心觀念
本題考查 Amdahl’s Law(阿姆達定律):程式整體的加速效果,受限於未被加速的部分。
若原程式中有比例 的執行時間可加速,該部分加速倍數為 ,則整體加速比為:
其中:
- 浮點運算比例:
- 浮點單元加速倍數:
- 非浮點運算比例:
解題方法
設原本程式的總執行時間為 。
原本浮點指令耗時:
非浮點指令耗時:
浮點單元加速 倍後,浮點指令的新耗時為:
非浮點指令未被加速,因此耗時仍為:
新的總執行時間:
第 Q3.1 題8 分
- Instruction Set Architecture (15 points)
Q3.1 (8 points)
A processor designer is considering adding complex instructions that directly operate on memory operands (thus potentially doing more work per instruction), which would reduce the number of instructions in programs but increase the complexity and latency of each instruction. Considering architectural trade-off principles, please explain why modern instruction set architectures favor simple, RISC-style instructions even if this means executing more instructions to accomplish the same task.
登入後即可作答並保存紀錄。
核心觀念
本題考查 RISC(Reduced Instruction Set Computer)與 CISC(Complex Instruction Set Computer)的架構取捨,以及處理器效能公式:
其中:
- Instruction Count(IC):程式實際執行的指令數量。
- CPI:平均每條指令所需的時脈週期數。
- :一個時脈週期的時間,與處理器時脈頻率互為倒數。
複雜指令可降低 IC,但也可能提高 CPI、增加單條指令延遲,並使處理器的關鍵路徑變長、時脈頻率下降。因此,不能只以「指令數較少」判斷效能。
解題方法
比較複雜指令與簡單指令對效能公式三個因素的影響:
1. 簡單指令有利於降低 CPI 與提高時脈頻率
RISC 通常採用下列設計:
- load/store 架構:只有載入與儲存指令存取記憶體,算術指令主要操作暫存器。
- 固定長度或規則化的指令格式。
- 較少的定址模式。
- 每條指令只完成單純且明確的工作。
- 指令延遲較容易預測。
例如:
LOAD R1, [A]
LOAD R2, [B]
ADD R3, R1, R2
STORE [C], R3
上述指令將記憶體存取、算術運算與結果寫回分開。這種規則化結構可使指令擷取、解碼、執行、記憶體存取與寫回階段形成深度管線,並降低管線控制的複雜度。
簡單指令不代表每條指令必定只需一個週期,而是代表硬體更容易安排固定且可預測的執行流程,從而降低平均 CPI 並提高時脈頻率。
2. 複雜記憶體運算指令會增加硬體負擔
例如下列指令一次完成位址計算、記憶體讀取、算術運算與結果寫回:
ADD R1, [R2 + offset]
處理器必須依序處理:
- 計算有效位址 。
- 存取資料快取。
- 取出記憶體資料。
- 執行加法。
- 將結果寫回暫存器。
這會使單條指令包含多個微操作,並造成下列問題:
- 指令解碼器與控制器更複雜。
- 管線階段難以保持規則。
- 記憶體延遲會直接影響算術指令。
- 快取未命中、TLB 未命中與頁面錯誤等例外處理更困難。
- 資料相依性、管線停滯與精確例外的維持更複雜。
- 指令的關鍵路徑變長,限制最高時脈頻率。
因此,雖然一條複雜指令能取代多條簡單指令,使 IC 下降,但其 CPI 與 可能大幅上升。
3. 簡單指令更適合現代平行化處理器
現代處理器通常使用:
- 超純量發射(superscalar issue)
- 亂序執行(out-of-order execution)
- 暫存器重新命名(register renaming)
- 分支預測
- 管線化
第 Q3.2 題7 分
Q3.2 (7 points)
In modern 64-bit processors with fixed-length instructions and deep pipelines, conditional branch instructions almost always use PC-relative addressing, where the branch target is encoded as an offset from the current program counter, instead of using an absolute target address.
Consider the following design and system context:
• Programs may be loaded at different memory addresses due to dynamic linking, shared libraries, or address space layout randomization.
• Branch instructions are frequent and often target nearby instructions such as loop back-edges or conditional blocks.
• Deep pipelines require early branch target calculation to reduce control hazards.
• Instruction cache performance is sensitive to instruction size and code density.
Based on the above context, please answer the following two questions:
(a) (4 points) State two reasons at the instruction-set / system level why PC-relative addressing is preferred over absolute addressing.
(b) (3 points) State one reason at the microarchitectural level why PC-relative addressing is especially beneficial for pipelined processors.
Your answers should be concise and focus on architectural principles.
登入後即可作答並保存紀錄。
核心觀念
PC-relative addressing 不直接儲存完整目標位址,而是儲存「目前 PC 與分支目標之間的位移量」:
其中 為分支目標位址, 為指令中編碼的帶符號 offset。若指令以固定大小對齊,offset 也可能先左移數位,例如:
PC_base 依 ISA 定義,可能是目前指令的 PC 或下一條指令的 PC。
解題方法
本題分成兩個層次:
- 指令集/系統層級:觀察程式搬移與指令編碼大小的影響。
- 微架構層級:觀察處理器如何在 pipeline 早期算出分支目標,並及早改變取指位址。
(a) 指令集/系統層級的兩個理由
1. 支援位置獨立程式碼與程式搬移
假設程式整體搬移 個位址單位,原本的 PC 與目標位址分別為 、,搬移後變成:
因此新的位移量為:
可見 PC-relative offset 不因程式載入位置改變而改變。程式因此能配合:
- 動態連結與共享函式庫;
- Address Space Layout Randomization(ASLR);
- 不同的載入位址與記憶體配置。
相較之下,absolute addressing 必須在載入時修改指令中的完整目標位址,產生 relocation 成本,也降低程式碼頁面共享的便利性。
2. 指令較精簡,提升程式碼密度與 instruction cache 效能
分支通常跳往附近的指令,例如迴圈回跳或條件區塊。這些目標只需要表示一個較小的 signed offset,不必在指令中放入完整的 64-bit absolute address。
第 Q4.1 題8 分
- Computer Arithmetic (15 points)
Q4.1 (8 points)
A processor represents signed integers using 8-bit two's complement format. The most significant bit is used as the sign bit, and arithmetic is performed using standard two's complement rules. Consider the following design and usage context:
• All arithmetic operations are performed modulo .
• Negation is implemented by bitwise inversion followed by adding 1.
• The same hardware is used for both signed and unsigned arithmetic.
Please answer the following three questions:
(a) (3 points) determine the minimum and maximum integers that can be represented in this system.
(b) (3 points) Explain why the set of representable values is asymmetric, i.e., why there is one more negative value than positive value.
(c) (2 points) Identify which specific bit pattern causes this asymmetry and explain its role in two's complement arithmetic.
Your answers should be concise and based on representation properties rather than examples alone.
登入後即可作答並保存紀錄。
核心觀念
8 位元二補數的數值定義為
其中 為符號位:
- :表示非負數。
- :最高位貢獻 。
總共有 種位元組合,且所有運算皆以模 計算。
解題方法
(a) 可表示的最小值與最大值
當符號位為 時,其餘 7 位可表示
因此最大值為
當符號位為 時,數值範圍為
即
因此最小值為
整體可表示範圍為
也就是
(b) 為何負數比正數多一個
符號位為 的 128 種組合表示
符號位為 的 128 種組合表示
因此負數共有 128 個,正數只有 127 個;多出的 1 個位元組合用來表示 。
根本原因是二補數必須保留全零位元組合表示 ,而零只有一種表示法。正負數配對後,正數範圍只能到 ,負數則可延伸至 ,形成
這個不對稱範圍。
(c) 造成不對稱的位元組合
造成不對稱的特殊位元模式是
依二補數定義,其數值為
第 Q4.2 題7 分
Q4.2 (7 points)
A processor implements floating-point arithmetic using a fixed-precision binary floating-point format that rounds the result of every arithmetic operation to the nearest representable value. Consider the following scenario:
You are given a very large array of floating-point numbers whose magnitudes vary widely.
Two programs compute the sum of all elements in the array:
• Program A sums the elements from left to right.
• Program B sorts the elements by magnitude and sums from smallest to largest.
Both programs use the same floating-point format and run on the same hardware. Despite using identical arithmetic instructions, the two programs produce different results.
Please answer the following questions:
(a) (4 points) State two fundamental properties of floating-point arithmetic that explain why the two programs may produce different sums.
(b) (3 points) Explain which summation order is expected to produce a more accurate result and why.
Your answers should refer to representation and arithmetic properties of floating-point numbers, not implementation bugs or software errors.
登入後即可作答並保存紀錄。
核心觀念
本題考查浮點數的有限精度、捨入誤差,以及浮點加法不具結合律。
具有 位有效二進位數字的浮點數,只能表示離散集合中的數值。對一般正常數而言,精確結果經過就近捨入可表示為
其中 為單位捨入誤差;二進位格式下通常有 。因此,實數的精確加法結果不一定能被格式精確表示。
此外,浮點數在不同數值尺度的間距不同。數值越大,相鄰可表示數值之間的絕對距離越大;若小數值小於大數值附近約半個 ULP,加入後可能完全被捨去,稱為吸收現象。
解題方法
令 表示將精確和 捨入至最近的可表示浮點數。
程式 A 的計算形式為
程式 B 先將資料依大小排列為
再計算
在實數精確運算中,交換加法順序不會改變總和;但浮點運算每一步都會捨入,因此中間結果不同,後續捨入結果也會不同。
(a)兩項基本性質
-
浮點數是有限精度表示,且每次運算都會產生捨入
有限有效位數使得許多精確結果無法直接表示,必須取最近的可表示值。每一次加法的誤差會成為下一次運算的輸入,累積後便可能造成不同的最終結果。
浮點數的可表示間距也會隨指數改變。當目前的部分和很大、待加入的數很小時,小數值可能落在相鄰可表示值的捨入範圍內而消失:
這種情形稱為小數值被大數值吸收。
-
浮點加法不具結合律
一般而言,
程式 A 與程式 B雖然使用相同的加法指令,但每次加法的操作數與中間部分和不同,相當於採用了不同的運算順序,因此捨入誤差也不同。
(b)較準確的加總順序
在題目的標準情境——陣列元素為非負數或大致同號——程式 B 由小到大加總,通常會得到較準確的結果。
原因是:
第 Q5.1 題8 分
- Pipelining & Hazards (15 points)
Q5.1 (8 points)
Consider a 5-stage pipeline (IF, ID, EX, MEM, WB). A load instruction is immediately followed by an instruction that uses the loaded value (for example, LW R1, 0(R2) followed by ADD R3, R1, R4). Please explain in detail: (1) (3 points) Why a data hazard occurs in this scenario, (2) (3 points) Why forwarding alone is insufficient to resolve it, and (3) (2 points) The minimum pipeline stall required (if any) to handle the hazard.
登入後即可作答並保存紀錄。
核心觀念
本題考查五級管線中的:
- 資料相依(data dependence)
- RAW(Read After Write,先寫後讀)資料危障
- Forwarding/bypassing
- Load-use hazard
- Pipeline stall/bubble
指令如下:
LW R1, 0(R2)
ADD R3, R1, R4
第一條 LW 是生產者(producer),會產生 R1 的新值;第二條 ADD 是消費者(consumer),必須讀取 R1 的新值。
(1)為何會發生資料危障
LW 的執行流程為:
- 在
EX階段計算記憶體位址; - 在
MEM階段讀取資料; - 在
WB階段將資料寫回R1。
因此,R1 的新值直到 LW 的 MEM 階段結束後才產生。
未加入停頓時,兩條指令的管線時序如下:
| Cycle | LW R1, 0(R2) | ADD R3, R1, R4 |
|---|---|---|
| 1 | IF | |
| 2 | ID | IF |
| 3 | EX | ID |
| 4 | MEM | EX |
| 5 | WB | MEM |
| 6 | WB |
ADD 在第 4 個 cycle 的 EX 階段便需要使用 R1,但 LW 必須到第 4 個 cycle 的 MEM 階段結束才取得記憶體資料。
所以,ADD 所需的 R1 新值尚未產生;若直接執行,ADD 會取得舊的 R1 值,造成錯誤。
這是典型的 RAW 資料危障:
雖然兩條指令的暫存器讀寫順序符合程式語意,但實際產生資料的時間晚於下一條指令使用資料的時間,因此形成管線危障。
(2)為何單靠 forwarding 不足以解決
Forwarding 的作用,是將尚未寫回暫存器檔案的結果,直接從後段管線轉送至使用該結果的指令。
例如:
ADD R1, R2, R3
SUB R4, R1, R5
第一條 ADD 的 ALU 結果在 EX 階段結束時已經產生;第二條 SUB 下一個 cycle 進入 EX 時,可以透過 EX/MEM → EX forwarding 取得結果,因此不必停頓。
但 LW 不同於一般 ALU 指令:
LW的EX階段只計算記憶體位址;- 真正載入的資料要到
MEM階段結束才產生; - 緊接在後的
ADD卻在同一個 cycle 進入EX。
未停頓時,第 4 個 cycle 的狀態是:
| 指令 | 所在階段 |
|---|---|
LW | MEM |
ADD | EX |
此時:
第 Q5.2 題7 分
Q5.2 (7 points)
Modern high-performance processors often use deep pipelines to increase clock frequency. However, deeper pipelines suffer higher penalties when branch predictions are incorrect. Assume the following simplified model:
• A branch is resolved at the end of the pipeline
• A mis-predicted branch causes all younger instructions in the pipeline to be flushed
• Branch prediction accuracy is fixed at 90%
• Each pipeline stage takes 1 cycle
Please answer the following two parts.
(a) Numerical Analysis (3 points)
For the following pipeline depths, please compute the branch misprediction penalty and the average cost per branch. Please draw a new table on your answer sheet for your answers.
| Pipeline Depth (stages) | Misprediction Penalty (cycles) | Average Cost per Branch (cycles) |
|---|---|---|
| 5 | ? | ? |
| 10 | ? | ? |
| 20 | ? | ? |
(b) Based on your numerical results and graph, explain why deeper pipelines rely more heavily on accurate branch prediction, even when prediction accuracy does not change. (4 points)
登入後即可作答並保存紀錄。
核心觀念
設管線深度為 stages。
本題採用簡化模型:分支指令須經過整條管線,直到最後一級才解析;若預測錯誤,錯誤路徑所造成的管線延遲視為 cycles。因此:
分支預測正確率為 ,故誤判率為:
平均每個分支因預測造成的額外成本為:
正確預測時不產生額外 penalty。
(a) 數值分析
當 時:
當 時:
當 時:
| 管線深度(stages) | 誤判懲罰(cycles) | 平均每分支成本(cycles) |
|---|---|---|
| 5 | 5 | 0.5 |
| 10 | 10 | 1.0 |
| 20 | 20 | 2.0 |
(b) 深管線為何更依賴準確的分支預測
平均成本與管線深度的關係為:
因此圖形是一條通過原點、斜率為 的直線:
第 Q6.1 題8 分
- Cache & Memory Hierarchy (15 points)
Q6.1 (8 points)
A system has a cache and memory with the following characteristics:
• Cache hit time = 1 cycle
• Cache miss rate = 2% (0.02)
• Miss penalty (time to fetch from lower-level memory on a miss) = 100 cycles
Please calculate the average memory access time (AMAT) for this system.
登入後即可作答並保存紀錄。
核心觀念
平均記憶體存取時間(Average Memory Access Time, AMAT)包含兩部分:
- 每次存取都必須付出的 Cache hit time。
- 只有 Cache miss 才會額外付出的 miss penalty。
標準公式為:
解題方法
代入題目數值:
亦可用 100 次存取驗算:
第 Q6.2 題7 分
Q6.2 (7 points)
Increasing cache associativity reduces conflict misses but may degrade performance. Please explain why increasing associativity eventually hurts performance, using architectural cost reasoning.
登入後即可作答並保存紀錄。
核心觀念
本題考查快取的「組相聯度」(set associativity)與平均記憶體存取時間(AMAT)之間的取捨。
若每個 set 有 個 ways,則:
- :direct-mapped cache。
- :2-way set-associative cache。
- 等於快取總 cache lines 數:fully associative cache。
增加 可以使同一個 set 中的資料有更多位置可放置,因此降低 conflict miss。但它不會消除:
- compulsory miss:第一次存取資料所造成的 miss;
- capacity miss:快取容量不足所造成的 miss。
因此,當 associativity 已經提高到一定程度後,繼續增加 所能減少的 miss 數量會快速遞減。
解題方法:利用 AMAT 分析
平均記憶體存取時間可表示為:
其中:
- :快取命中時間;
- :miss rate;
- :miss penalty。
增加 associativity 時:
- 通常下降,這是降低 conflict miss 的好處。
- 通常上升,這是硬體成本造成的代價。
比較兩種 associativity 與 :
當 associativity 很高時, 會變得很小。若命中時間增加量滿足:
則:
此時增加 associativity 反而使整體效能下降。
增加 associativity 的硬體成本
以一次 cache access 為例,處理器必須根據 address 分出:
- block offset;
- set index;
- tag。
在 -way set-associative cache 中,同一個 set 內的 個 tag 通常需要同時比較:
- 讀取 個 tag;
- 使用 組 tag comparator 與要求的 tag 比較;
- 判斷哪一個 way 命中;
- 使用多工器(multiplexer)選出對應的 data;
- 更新 replacement policy,例如 LRU 或 pseudo-LRU 資訊。
因此,增加 會造成下列影響: