113 年 國立中山大學資訊工程學系碩士班甲組《計算機結構》

📄 試題原卷 免費註冊後即可對照原始考卷 PDF免費註冊

第 1 題8 分

(8% total) When parallelizing an application, the ideal speedup is speeding up by the number of processors. This is limited by two things: percentage of the application that can be parallelized and the cost of communication. If 80% of the application is parallelizable, answer the following questions.
1.1 (4%) What is the speedup with 8 processors, ignoring the cost of communication?
1.2 (4%) What is the speedup with 8 processors if, for every time the number of processors is doubled, the communication overhead is increased by 0.5% of the original execution time?

登入後即可作答並保存紀錄。

這一題的完整詳解

核心觀念

本題的核心考點為平行計算架構中的阿姆達爾定律(Amdahl's Law)以及通訊開銷(Communication Overhead)對平行加速效能的影響。

  1. 阿姆達爾定律(Amdahl's Law):
    定義加速比(Speedup)為單一處理器之執行時間(T1T_1)與 pp 個處理器平行執行時間(TpT_p)的比值:
    Speedup=ToriginalTparallel=T1Tp\text{Speedup} = \frac{T_{\text{original}}}{T_{\text{parallel}}} = \frac{T_1}{T_p}
    若程式中可平行化比例為 ff,不可平行化(循序執行)比例為 1−f1 - f,在忽略通訊與同步開銷的情況下:
    Tp=T1×((1−f)+fp)T_p = T_1 \times \left( (1 - f) + \frac{f}{p} \right)
    Speedup=1(1−f)+fp\text{Speedup} = \frac{1}{(1 - f) + \dfrac{f}{p}}

  2. 納入通訊開銷之平行執行時間:
    當考量通訊額外開銷(TcommT_{\text{comm}})時,總平行執行時間變為:
    Tp=T1×((1−f)+fp)+TcommT_p = T_1 \times \left( (1 - f) + \frac{f}{p} \right) + T_{\text{comm}}
    加速比則修正為:
    Speedup=T1Tp=1(1−f)+fp+TcommT1\text{Speedup} = \frac{T_1}{T_p} = \frac{1}{(1 - f) + \dfrac{f}{p} + \dfrac{T_{\text{comm}}}{T_1}}


解題方法

令單一處理器的原始執行時間為基準 T1=1T_1 = 1。
已知可平行化比例 f=80%=0.8f = 80\% = 0.8,循序比例 1−f=20%=0.21 - f = 20\% = 0.2,處理器個數 p=8p = 8。

1.1 計算忽略通訊成本時的加速比(p=8p = 8)

  1. 計算平行執行時間 T8T_8:
    T8=T1×((1−f)+f8)=T1×(0.2+0.88)=T1×(0.2+0.1)=0.3×T1T_8 = T_1 \times \left( (1 - f) + \frac{f}{8} \right) = T_1 \times \left( 0.2 + \frac{0.8}{8} \right) = T_1 \times (0.2 + 0.1) = 0.3 \times T_1
🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 2 題8 分

(8% total) Assume a hypothetical GPU with the following characteristics:
■ Clock rate 2.0 GHz
■ Contains 16 SIMD processors, each containing 16 single-precision floating-point units
■ Has 600 GB/sec off-chip memory bandwidth
2.1 (4%) Without considering memory bandwidth, what is the peak single-precision floating-point throughput for this GPU in GFLOPS/sec, assuming that all memory latencies can be hidden?
2.2 (4%) Is this throughput sustainable given the memory bandwidth limitation? Please explain your reasons clearly.

登入後即可作答並保存紀錄。

這一題的完整詳解

核心觀念

本題考查兩個概念:

  1. GPU 的峰值浮點吞吐量
    計算所有可同時工作的浮點單元數,再乘上時脈頻率。

  2. 記憶體頻寬限制與算術強度(Arithmetic Intensity)
    算術強度定義為:

    AI=浮點運算次數從 off-chip memory 搬移的位元組數AI=\frac{\text{浮點運算次數}}{\text{從 off-chip memory 搬移的位元組數}}

    GPU 的實際效能受下式限制:

    Pactual≤min⁡(Pcompute, Bmemory×AI)P_{\text{actual}}\leq \min\left(P_{\text{compute}},\ B_{\text{memory}}\times AI\right)

    即使記憶體延遲能完全隱藏,有限的記憶體頻寬仍可能限制吞吐量。

解題方法

2.1 峰值單精度浮點吞吐量

GPU 中的單精度浮點單元總數為:

16 個 SIMD processors×16 個浮點單元=256 個浮點單元16\ \text{個 SIMD processors}\times 16\ \text{個浮點單元}=256\ \text{個浮點單元}

時脈頻率為:

2.0 GHz=2.0×109 cycles/sec2.0\ \text{GHz}=2.0\times 10^9\ \text{cycles/sec}

在每個浮點單元每個 clock 完成 1 次浮點運算的假設下:

Pcompute=256×2.0×109=512×109 FLOP/secP_{\text{compute}} =256\times 2.0\times 10^9 =512\times 10^9\ \text{FLOP/sec}

因此:

Pcompute=512 GFLOPSP_{\text{compute}}=512\ \text{GFLOPS}

2.2 記憶體頻寬是否足以維持此吞吐量

要維持 512 GFLOPS512\ \text{GFLOPS},所需的最低算術強度為:

AIbreak-even=512 GFLOPS600 GB/sec=512600≈0.853 FLOP/byteAI_{\text{break-even}} =\frac{512\ \text{GFLOPS}}{600\ \text{GB/sec}} =\frac{512}{600} \approx 0.853\ \text{FLOP/byte}

因此:

  • 若 AI≥0.853 FLOP/byteAI\geq 0.853\ \text{FLOP/byte},記憶體頻寬足以支撐計算單元達到 512 GFLOPS512\ \text{GFLOPS}。
  • 若 AI<0.853 FLOP/byteAI<0.853\ \text{FLOP/byte},效能會受記憶體頻寬限制,無法達到峰值。

題目未提供實際程式的資料搬移量,因此嚴格而言缺少算術強度,無法對所有程式給出唯一的「是」或「否」。

🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 3 題18 分

(18% total) Assume that individual stages of the datapath have the following latencies, and instructions executed by the processor are broken down as the following percentages:
IF
ID
EX
MEM
WB
250 ps
150 ps
350 ps
300 ps
200 ps
alu
45%
beq
lw
SW
10%
25%
20%
3.1 (4%) What is the clock cycle time in a non-pipelined and pipelined processor, respectively?
3.2 (4%) What is the individual total latency of the following two instructions: lw, beq in a non-pipelined and pipelined processor?
3.3 (6%) Assuming there are no stalls or hazards, what is the total execution time of the following four instructions: lw, sw, add, beq in a non-pipelined and pipelined processor, respectively?
3.4 (4%) Assuming there are no stalls or hazards, what is the utilization of the data memory and the write-register port of the "Register" unit?

登入後即可作答並保存紀錄。

這一題的完整詳解

3.1 時脈週期時間 (Clock Cycle Time)

  1. 非流水線處理器 (Non-pipelined / Single-cycle):
    時脈週期時間由執行時間最長的指令(即經過全部 5 個階段的 lw 指令)決定:
    Cycle Time=IF+ID+EX+MEM+WB=250+150+350+300+200=1250 ps\text{Cycle Time} = \text{IF} + \text{ID} + \text{EX} + \text{MEM} + \text{WB} = 250 + 150 + 350 + 300 + 200 = 1250\text{ ps}

  2. 流水線處理器 (Pipelined):
    時脈週期時間由延遲最長的單一階段(EX 階段)決定:
    Cycle Time=max⁡(IF,ID,EX,MEM,WB)=350 ps\text{Cycle Time} = \max(\text{IF}, \text{ID}, \text{EX}, \text{MEM}, \text{WB}) = 350\text{ ps}

【答案】非流水線處理器為 1250 ps1250\text{ ps};流水線處理器為 350 ps350\text{ ps}。


3.2 個別指令總延遲 (Individual Total Latency)

  1. 非流水線處理器:
    單一時脈週期架構下,時脈週期固定為 1250 ps1250\text{ ps},故每筆指令執行時間皆佔據一個完整時脈週期 1250 ps1250\text{ ps}(若按硬體實際經過的階段延遲,lw 為 1250 ps1250\text{ ps},beq 為 IF+ID+EX=250+150+350=750 ps\text{IF} + \text{ID} + \text{EX} = 250 + 150 + 350 = 750\text{ ps})。

  2. 流水線處理器:
    所有指令在流水線中皆須通過 5 個 pipeline 階段,總延遲為:
    Latency=5×350 ps=1750 ps\text{Latency} = 5 \times 350\text{ ps} = 1750\text{ ps}
    因此 lw 與 beq 的總延遲皆為 1750 ps1750\text{ ps}。

🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 4 題15 分

(15% total) The importance of having a good branch predictor depends on how often conditional branches are executed. Together with branch predictor accuracy, this will determine how much time is spent stalling due to mispredicted branches. Assume that the breakdown of dynamic instructions into various instruction categories is as follows:
R-Type
40%
BEQ
20%
JMP
5%
LW
25%
SW
10%
Also, assume the following branch predictor accuracies:
Always-Taken
40%
Always-Not-Taken
60%
2-Bit
80%
4.1 (5%) Assume no stall cycle is required when branch prediction is correct. Stall cycles due to mispredicted branches increase the CPI. What is the extra CPI due to mispredicted branches with the always-taken predictor? Assume that branch outcomes are determined in the EX stage of a basic 5-stage pipelined processor, that there are no data hazards, and that no delay slots are used.
4.2 (5%) Repeat 4.1 for the 2-bit predictor.
4.3 (5%) With the always-taken predictor, what speedup would be achieved if we could convert half of the branch instructions in a way that replaces a branch instruction with an ALU instruction? Assume that correctly and incorrectly predicted instructions have the same chance of being replaced.

登入後即可作答並保存紀錄。

這一題的完整詳解

觀念與前提設定

  • 條件分歧指令(BEQ)比例:20%20\%
  • 分支預測錯誤懲罰(Penalty):在傳統 5 階段 Pipeline(IF, ID, EX, MEM, WB)中,若分支結果於 EX 階段(第 3 階段)確定,預測錯誤時需 Flush 已進入 IF 與 ID 階段的 2 個指令,懲罰時間為 22 個 Clock Cycles。
  • 基礎 CPI:無任何 Hazard 下之理想 CPI 為 1.01.0。

4.1

使用 Always-Taken 預測器,正確率為 40%40\%,預測錯誤率(Misprediction Rate)為 1−40%=60%1 - 40\% = 60\%。

預測錯誤導致的額外 CPI(Extra CPI)計算如下:
Extra CPI=BEQ Frequency×Misprediction Rate×Penalty Cycles\text{Extra CPI} = \text{BEQ Frequency} \times \text{Misprediction Rate} \times \text{Penalty Cycles}
Extra CPI=0.20×(1−0.40)×2=0.20×0.60×2=0.24\text{Extra CPI} = 0.20 \times (1 - 0.40) \times 2 = 0.20 \times 0.60 \times 2 = 0.24

【答案】0.240.24


4.2

使用 2-Bit 預測器,正確率為 80%80\%,預測錯誤率為 1−80%=20%1 - 80\% = 20\%。

🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 5 題15 分

(15% total) Caches are important to providing a high-performance memory hierarchy to processors. Below is a list of 32-bit memory address references, given as word addresses.
3, 180, 43, 2, 191, 88, 190, 14, 181, 44, 186, 253
5.1 (9%) For each of these references, identify the binary address, the tag, and the index given a direct-mapped cache with two-word blocks and a total size of 8 blocks. Also list if each reference is a hit or a miss, assuming the cache is initially empty.
5.2 (6%) How many total bits are required for a direct-mapped cache with 32 KiB of data, 8-word blocks, and 32-bit address?

登入後即可作答並保存紀錄。

這一題的完整詳解

5.1 記憶體存取紀錄分析

快取記憶體位址分割(32-bit 位址)

  • 單一區塊大小:2 words=8 bytes⇒2\text{ words} = 8\text{ bytes} \Rightarrow Offset 占 3 bits3\text{ bits}(含 1 bit1\text{ bit} Block Offset 與 2 bits2\text{ bits} Byte Offset)。
  • 區塊總數:8 blocks⇒8\text{ blocks} \Rightarrow Index 占 log⁡2(8)=3 bits\log_2(8) = 3\text{ bits}。
  • Tag 長度:32−3−3=26 bits32 - 3 - 3 = 26\text{ bits}。

位址切割結構如下:

  • Tag=Word Address≫4\text{Tag} = \text{Word Address} \gg 4(前 26 位元)
  • Index=(Word Address≫1)(mod8)\text{Index} = (\text{Word Address} \gg 1) \pmod 8(中間 3 位元)
  • Block Offset=Word Address(mod2)\text{Block Offset} = \text{Word Address} \pmod 2(最低 1 位元)

各筆記憶體存取結果推導

  1. Word Address = 3 (00000000000000000000000000000011200000000000000000000000000000011_2)

    • 32-bit 位址:0x0000000C (0000…000000110020000\dots0000001100_2)
    • Tag: 00 (0…0020\dots00_2),Index: 11 (0012001_2)
    • 狀態:Miss(將 Block 1 填入 Tag 0)
  2. Word Address = 180 (00000000000000000000000010110100200000000000000000000000010110100_2)

    • 32-bit 位址:0x000002D0 (0000…00101101000020000\dots001011010000_2)
    • Tag: 1111 (0…0101120\dots01011_2),Index: 22 (0102010_2)
    • 狀態:Miss(將 Block 2 填入 Tag 11)
  3. Word Address = 43 (00000000000000000000000000101011200000000000000000000000000101011_2)

    • 32-bit 位址:0x000000AC (0000…00001010110020000\dots000010101100_2)
    • Tag: 22 (0…0001020\dots00010_2),Index: 55 (1012101_2)
    • 狀態:Miss(將 Block 5 填入 Tag 2)
  4. Word Address = 2 (00000000000000000000000000000010200000000000000000000000000000010_2)

    • 32-bit 位址:0x00000008 (0000…000000100020000\dots0000001000_2)
    • Tag: 00 (0…0020\dots00_2),Index: 11 (0012001_2)
    • 狀態:Hit(Block 1 已有 Tag 0)
  5. Word Address = 191 (00000000000000000000000010111111200000000000000000000000010111111_2)

    • 32-bit 位址:0x000002FC (0000…00101111110020000\dots001011111100_2)
    • Tag: 1111 (0…0101120\dots01011_2),Index: 77 (1112111_2)
    • 狀態:Miss(將 Block 7 填入 Tag 11)
  6. Word Address = 88 (00000000000000000000000001011000200000000000000000000000001011000_2)

    • 32-bit 位址:0x00000160 (0000…00010110000020000\dots000101100000_2)
    • Tag: 55 (0…0010120\dots00101_2),Index: 44 (1002100_2)
    • 狀態:Miss(將 Block 4 填入 Tag 5)
  7. Word Address = 190 (00000000000000000000000010111110200000000000000000000000010111110_2)

    • 32-bit 位址:0x000002F8 (0000…00101111100020000\dots001011111000_2)
    • Tag: 1111 (0…0101120\dots01011_2),Index: 77 (1112111_2)
    • 狀態:Hit(Block 7 已有 Tag 11)
  8. Word Address = 14 (00000000000000000000000000001110200000000000000000000000000001110_2)

    • 32-bit 位址:0x00000038 (0000…00000011100020000\dots000000111000_2)
    • Tag: 00 (0…0020\dots00_2),Index: 77 (1112111_2)
    • 狀態:Miss(替換 Block 7 的 Tag 為 0)
  9. Word Address = 181 (00000000000000000000000010110101200000000000000000000000010110101_2)

    • 32-bit 位址:0x000002D4 (0000…00101101010020000\dots001011010100_2)
    • Tag: 1111 (0…0101120\dots01011_2),Index: 22 (0102010_2)
    • 狀態:Hit(Block 2 已有 Tag 11)
  10. Word Address = 44 (00000000000000000000000000101100200000000000000000000000000101100_2)

    • 32-bit 位址:0x000000B0 (0000…00001011000020000\dots000010110000_2)
🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 6 題16 分

(16% total) Assume that the CPI with a perfect cache is 2.0, the clock cycle time is 1.0 ns, there are 1.5 memory references per instruction, the size of both caches is 64 KB, and both have a block size of 64 bytes. One cache is direct mapped and the other is two-way set associative. Since the speed of the CPU is tied directly to the speed of a cache hit, assume the CPU clock cycle time must be stretched 1.3 times to accommodate the selection multiplexor of the set-associative cache. To the first approximation, the cache miss penalty is 80 ns for either cache organization. Assume the hit time is 1 clock cycle, the miss rate of a direct-mapped 64 KB cache is 1.5%, and the miss rate for a two-way set-associative cache of the same size is 1.1%.
6.1 (8%) Calculate the average memory access time for these two different cache organizations.
6.2 (8%) Calculate the CPU performance in terms of CPU time.

登入後即可作答並保存紀錄。

這一題的完整詳解

核心觀念

本題比較兩種快取組織的兩項指標:

  1. 平均記憶體存取時間:
    AMAT=Hit Time+Miss Rate×Miss Penalty\text{AMAT}=\text{Hit Time}+\text{Miss Rate}\times\text{Miss Penalty}

  2. CPU 執行時間:
    TCPU=Instruction Count×CPI×Clock Cycle TimeT_{\text{CPU}}=\text{Instruction Count}\times\text{CPI}\times\text{Clock Cycle Time}

完美快取下的 CPI=2.0\text{CPI}=2.0 已包含正常快取命中的執行成本,因此只需額外加入快取未命中的停滯時間。

快取總區塊數為:

64 KB64 B=1024 blocks\frac{64\text{ KB}}{64\text{ B}}=1024\text{ blocks}

直接映射快取有 10241024 條快取線;二路組合相聯快取有 512512 個集合,每集合含 22 條快取線。題目已直接給出兩者的 miss rate,因此不必再由位址存取序列推導 miss rate。


6.1 平均記憶體存取時間

直接映射快取

原始時脈週期為:

Tc=1.0 nsT_c=1.0\text{ ns}

命中時間為一個時脈週期:

Thit=1.0 nsT_{\text{hit}}=1.0\text{ ns}

因此:

AMATdirect=Thit+Miss Rate×Miss Penalty=1.0+(0.015)(80)=2.20 ns\begin{aligned} \text{AMAT}_{\text{direct}} &=T_{\text{hit}}+\text{Miss Rate}\times\text{Miss Penalty}\\ &=1.0+(0.015)(80)\\ &=2.20\text{ ns} \end{aligned}

二路組合相聯快取

由於選擇多工器使 CPU 時脈週期延長為原本的 1.31.3 倍:

Tc=1.3×1.0=1.3 nsT_c=1.3\times1.0=1.3\text{ ns}

命中時間為:

Thit=1.3 nsT_{\text{hit}}=1.3\text{ ns}

因此:

AMAT2-way=1.3+(0.011)(80)=1.3+0.88=2.18 ns\begin{aligned} \text{AMAT}_{\text{2-way}} &=1.3+(0.011)(80)\\ &=1.3+0.88\\ &=2.18\text{ ns} \end{aligned}

二路組合相聯快取的 AMAT 為 2.18 ns2.18\text{ ns},略低於直接映射快取的 2.20 ns2.20\text{ ns}。


6.2 CPU 執行時間

設程式共有 II 條指令。每條指令平均有 1.51.5 次記憶體參考,因此每條指令的 miss 次數為:

Misses per Instruction=1.5×Miss Rate\text{Misses per Instruction}=1.5\times\text{Miss Rate}

每條指令的額外 miss 停滯時間為:

額外時間=(1.5)(Miss Rate)(80 ns)\text{額外時間}=(1.5)(\text{Miss Rate})(80\text{ ns})

直接映射快取

每條指令的 miss 次數:

1.5×0.015=0.02251.5\times0.015=0.0225

每條指令的 miss 停滯時間:

🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 7 題20 分

(20% total) Consider the virtual memory system.
7.1 (8%) The following list provides parameters of a virtual memory system.
Virtual Address (bits)
42
Physical DRAM Installed
16 GiB
Page Size
4 KiB
PTE Size (byte)
4
For a single-level page table, how many page table entries (PTEs) are needed? How much physical memory is needed for storing the page table?
7.2 (12%) The following data constitutes a stream of virtual address as seen on a system. Assume 4 KiB pages, a 4-entry fully associative TLB, and true LRU replacement. If pages must be brought in from disk, increment the next largest page number.
TLB
Valid
Tag
Physical Page Number
1
11
12
1
7
4
1
3
6
0
4
9
Page table
Valid
Physical Page or in Disk
1
5
0
Disk
0
Disk
1
6
1
9
1
11
0
Disk
1
4
0
Disk
0
Disk
1
3
1
12
Given the following address stream: 4669, 13916, ..., and the initial TLB and page table states provided above, show the final state of the system. Also list for each reference if it is a hit in the TLB, a hit in the page table, or a page fault.

登入後即可作答並保存紀錄。

這一題的完整詳解

核心觀念

本題考查兩部分:

  1. 單層頁表的大小計算:

    • 虛擬位址由「虛擬頁號 VPN」與「頁內位移 offset」組成。
    • 頁表需要一個 PTE 對應一個虛擬頁。
    • 頁表大小=PTE 數量 × 每個 PTE 大小。
  2. TLB、頁表與 Page Fault 的判斷:

    • 先以虛擬位址查 TLB。
    • TLB 命中:直接取得 Physical Page Number。
    • TLB 未命中,再查頁表。
    • 頁表有效:稱為 page table hit,並將映射填入 TLB。
    • 頁表無效且標示 Disk:發生 page fault,從磁碟載入頁面,再更新頁表與 TLB。
    • TLB 採 4-entry fully associative,必須逐筆比對 Tag。
    • 替換策略為 true LRU,最近使用的項目保留,最久未使用的項目被替換。

7.1 單層頁表大小

步驟一:計算頁內位移位元數

頁面大小為 4 KiB4\text{ KiB}:

4 KiB=4×210=212 bytes4\text{ KiB}=4\times 2^{10}=2^{12}\text{ bytes}

因此頁內位移需要 1212 bits。

步驟二:計算虛擬頁號位元數

虛擬位址為 4242 bits,因此虛擬頁號為:

42−12=30 bits42-12=30\text{ bits}

可表示的虛擬頁數為:

230 pages2^{30}\text{ pages}

單層頁表每個虛擬頁需要一個 PTE,所以 PTE 數量為:

230 個 PTE\boxed{2^{30}\text{ 個 PTE}}

步驟三:計算頁表所需記憶體

每個 PTE 為 44 bytes:

230×4=230×22=232 bytes2^{30}\times 4 = 2^{30}\times 2^2 = 2^{32}\text{ bytes}

而:

232 bytes=4 GiB2^{32}\text{ bytes}=4\text{ GiB}

因此頁表需要:

4 GiB\boxed{4\text{ GiB}}

題目給出的 Physical DRAM Installed 為 1616 GiB,並不影響單層頁表本身的計算;它只表示系統實際安裝的實體記憶體容量。


7.2 位址轉換與 TLB 更新

已知系統

  • Page Size:4 KiB=40964\text{ KiB}=4096 bytes
  • TLB:4-entry fully associative
  • Replacement:true LRU
  • 虛擬頁號:
VPN=⌊Virtual Address4096⌋VPN=\left\lfloor \frac{\text{Virtual Address}}{4096}\right\rfloor
  • 頁內位移:
offset=Virtual Address mod 4096offset=\text{Virtual Address}\bmod 4096

初始 TLB

ValidTag(VPN)Physical Page Number
11112
174
136
049

其中 Tag 代表虛擬頁號 VPN。

初始 Page Table

VPNValidPhysical Page 或 Disk
015
10Disk
20Disk
316
419
5111
60Disk
714
80Disk
90Disk
1013
11112

題目中的完整位址串流只列出 4669, 13916, ...,省略號後的位址未提供,因此無法唯一決定原題要求的完整最終狀態。以下先依照目前明確列出的兩筆位址計算;此計算也展示完整串流的處理方式。


Reference 1:Virtual Address = 4669

計算 VPN 與 offset

4669=1×4096+5734669=1\times4096+573

因此:

VPN=1,offset=573VPN=1,\qquad offset=573

查詢 TLB

初始 TLB 的 Tag 為 11,7,311,7,3,沒有 VPN 11。

因此:

TLB miss\boxed{\text{TLB miss}}

查詢 Page Table

VPN 11 的頁表項為:

VPNValidPhysical Page 或 Disk
10Disk

該頁不在實體記憶體中,因此發生:

page fault\boxed{\text{page fault}}
🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

其他考古題