114 年 國立成功大學電機工程學系碩士班丁組《計算機組織》

📄 試題原卷 免費註冊後即可對照原始考卷 PDF免費註冊

第 1 題10 分

  1. (10pts, no partial point, no penalty) Abstractions is an important technique for complex design. Which of the following statements is/are TRUE?
    (a) Abstraction hides the complexities of hardware from software, allowing programmers to focus on high-level algorithms without needing to understand the underlying hardware specifics.
    (b) Abstraction refers to the process of directly controlling hardware components such as memory, registers, and buses in a low-level programming environment.
    (c) Abstraction in computer architecture is about optimizing hardware performance by reducing the number of components in the processor.
    (d) Abstraction eliminates the need for an instruction set architecture by directly controlling the execution of machine-level instructions.
    (e) Abstraction provides a clear separation between the hardware and software layers by defining an instruction set architecture (ISA) that allows software to execute on any hardware.

登入後即可作答並保存紀錄。

這一題的完整詳解

核心觀念

本題考查計算機組織與架構中的關鍵設計理念:利用抽象化簡化設計(Use Abstraction to Simplify Design)。

核心定義與架構原則:

  1. 抽象化(Abstraction):在計算機系統中,抽象化是指隱藏低階(Lower-level)物理與電路細節,僅對外提供簡化且規範化的高階(Higher-level)介面。其目的在於降低系統複雜度、提高軟體開發效率並促進軟硬體獨立演進。
  2. 指令集架構(Instruction Set Architecture, ISA):ISA 是計算機架構中最核心的抽象化層次(Hardware/Software Interface)。它定義了軟體(編譯器、作業系統、應用程式)能看見的硬體抽象模型(包含指令集、暫存器集合、記憶體定址模式等),使軟體開發者無須了解底層邏輯閘或管道(Pipelining)等具體微架構(Microarchitecture)細節。
  3. 軟硬體解耦與可攜性(Software Portability):透過 ISA 這一層抽象介面,軟體與硬體得以完全分離。只要硬體遵從該 ISA 規範,編譯後的二進制程式碼即可在任何實作該 ISA 的硬體實作上記錄與執行。

解題方法

本題為概念性多重選擇題,切入點為比對各選項敘述是否符合教科書(如 Patterson & Hennessy 的 Computer Organization and Design)中關於 Abstraction 與 ISA 的標準觀念:

  1. 檢驗選項是否正確表達「抽象化的目的」:遮蔽底層細節、提供高階模型。
  2. 檢驗選項是否混淆「抽象化」與「低階硬體控制/物理最佳化」:直接操作硬體屬於低階實作而非抽象化;減少處理器元件數量屬於硬體面積與成本最佳化。
  3. 檢驗選項是否正確認識「ISA 與抽象化的關係」:ISA 本身即為抽象化層次,絕非被抽象化所消除。
  4. 依據上述邏輯逐一分析各選項之真偽。

選項分析

  • (a) 正確。

    • 原文:Abstraction hides the complexities of hardware from software, allowing programmers to focus on high-level algorithms without needing to understand the underlying hardware specifics.
    • 解析:此敘述完全符合抽象化的基本定義。抽象化層次(如高階程式語言、作業系統、編譯器)將底層硬體(如電路結構、暫存器傳輸位準 RTL、匯流排時序)的複雜性隱藏起來,使程式設計師得以專注於演算法與高階邏輯開發,無須掌握低階硬體的物理實作細節。
  • (b) 錯誤。

    • 原文:Abstraction refers to the process of directly controlling hardware components such as memory, registers, and buses in a low-level programming environment.
🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 2 題10 分

  1. (10pts, no partial point, no penalty) In a shared-memory multiprocessor system that implements the MESI cache-coherence protocol, which of the following statements is/are TRUE?
    (a) The Modified (M) state indicates that the cache line is dirty and present only in the current cache.
    (b) The Shared (S) state indicates that the cache line may be present in multiple caches, but none of those copies is dirty.
    (c) The Exclusive (E) state indicates that the cache line is present in only one cache and is clean (not dirty).
    (d) The Invalid (I) state indicates that the cache line is not valid and cannot be used.
    (e) A cache line transitions directly from Modified (M) to Exclusive (E) when another cache attempts to read the same data.

登入後即可作答並保存紀錄。

這一題的完整詳解

核心觀念

本題考驗共享記憶體多處理器系統(Shared-memory Multiprocessor System)中經典的 MESI 快取一致性協定(Cache-Coherence Protocol) 之狀態定義與狀態轉移機制。

MESI 協定屬於基於窺探(Snooping-based)的寫回(Write-Back)快取一致性協定,其名稱由四種快取區段(Cache Line)狀態的首字母組成:

  1. M(Modified,已修改/修飾):快取區段僅存在於當前快取中(獨占),且內容已被修改,與主記憶體內容不一致(Dirty)。
  2. E(Exclusive,獨占):快取區段僅存在於當前快取中(獨占),且內容未被修改,與主記憶體內容一致(Clean)。
  3. S(Shared,共享):快取區段可能存在於多個快取中,且所有副本內容均未經修改,與主記憶體內容一致(Clean)。
  4. I(Invalid,無效):快取區段不包含有效資料,處理器無法直接使用。

解題方法

解題切入點分為兩個層面:

  1. 狀態定義比對:檢視 (a)~(d) 各選項對 M、S、E、I 定義的描述,對照 MESI 的「獨占性(Single/Multiple Copy)」與「一致性(Clean/Dirty)」兩大軸線判斷正確性。
  2. 狀態轉移分析:針對 (e) 選項,分析當其他快取發出讀取請求(BusRd)時,原持有 M 狀態快取區段的狀態轉移路徑。

選項分析

  • (a) 正確
    分析:Modified (M) 狀態代表該快取區段僅由當前快取獨家持有(Present only in current cache),且由於執行過寫入操作,資料為 Dirty 狀態,主記憶體中的資料已過期。

  • (b) 正確
    分析:Shared (S) 狀態表示該資料可能同時存在於多個處理器的快取中。

🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 3 題10 分

  1. (10pts, no partial point, no penalty) In a multiprocessor system, the memory consistency model defines the order in which operations (reads and writes) appear to execute from the perspective of different processors. Consider the Sequential Consistency (SC) model. Which of the following statements about Sequential Consistency is/are TRUE?
    (a) Under SC, all memory accesses appear to execute in a single global order that respects each processor's program order.
    (b) Under SC, a processor is free to reorder its loads and stores arbitrarily to improve performance.
    (c) SC is stronger (i.e., imposes more ordering constraints) than weaker consistency models, such as Release Consistency (which allowing certain memory operations to be reordered or overlapped).
    (d) Implementing SC often requires hardware or compiler mechanisms (such as memory fences) to ensure the correct ordering of memory operations.
    (e) Under SC, a read on one processor may return a value written by another processor before that write is visible to all other processors.

登入後即可作答並保存紀錄。

這一題的完整詳解

核心觀念

本題考查多處理器系統(Multiprocessor Systems)中的記憶體一致性模型(Memory Consistency Models),特別是**順序一致性模型(Sequential Consistency, SC)**的學理定義、約束強度、與弱一致性模型(Weak Consistency Models)的比較以及實作機制。

根據 Leslie Lamport 於 1979 年提出的定義,順序一致性(SC)必須嚴格滿足以下兩大核心要件:

  1. 程式順序(Program Order):單一處理器內部發出的所有記憶體存取操作(Reads 與 Writes),在全域執行順序中必須完全符合該處理器本身的程式順序。
  2. 全域單一順序與寫入原子性(Global Single Total Order & Write Atomicity):所有處理器的記憶體操作在邏輯上呈現為一個單一全域順序(Single Global Order)交錯執行;且任何寫入操作對所有處理器而言必須同時可見(Write Atomicity)。

解題方法

本題切入點為將選項敘述逐一對照 SC 的兩大核心要件以及微架構實作原理:

  1. 記憶體操作順序:判斷處理器或編譯器是否能對存取進行重排(Reordering)。
  2. 模型的約束強度:比較 SC 與弱一致性模型(如 Release Consistency)在存取限制上的強弱關係。
  3. 硬體/軟體實作代價:分析現代亂序執行(Out-of-Order Execution)處理器如何透過 Memory Fence 等機制強制維持 SC。
  4. 寫入可見性(Visibility):驗證寫入操作是否具備原子性(Multi-copy Atomicity)。

選項分析

  • (a) 正確。
    此敘述為 Lamport 對於順序一致性(SC)的標準定義。在 SC 規範下,多個處理器發出的所有 Load 與 Store 操作,在邏輯上均可交錯排列成一個統一的全域執行順序(Single Global Total Order),且該順序完全尊重各個處理器內部的程式順序(Program Order)。
🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 4 題10 分

  1. (10pts, no partial point, no penalty) Modern superscalar processors use out-of-order execution, branch prediction, and speculative execution to improve performance. Which of the following statements is/are TRUE?
    (a) Branch prediction reduces pipeline stalls by anticipating the outcome of a branch.
    (b) Out-of-order execution allows the processor to execute independent instructions while waiting for data dependencies to resolve.
    (c) Out-of-order execution can reorder instructions with true data dependencies (RAW) without any extra hardware or compiler support.
    (d) Precise exceptions require instructions to complete in-order at the commit stage, even if they are executed out-of-order.
    (e) When a branch is mispredicted, only the instructions that have already committed must be flushed from the pipeline.

登入後即可作答並保存紀錄。

這一題的完整詳解

核心觀念

本題考驗現代**超純量亂序執行處理器(Superscalar Out-of-Order Processor)**的核心微架構技術,涵蓋以下四大關鍵觀念:

  1. 分支預測(Branch Prediction)與控制冒險(Control Hazard):透過在取指階段(Fetch Stage)提前預測分支方向與目標位址,避免流水線產生停頓(Stalls / Bubbles)。
  2. 亂序執行(Out-of-Order Execution, OoO)與資料相依性(Data Dependencies):
    • 真相依性(True Data Dependency / Flow Dependency, RAW - Read After Write):代表資料的實質傳遞關係,後續指令必須等待前導指令產出數值,不可任意重排提前執行。
    • 偽相依性(False Dependencies, WAR / WAW):僅為暫存器名稱衝突,可透過暫存器重命名(Register Renaming)解除。
    • 亂序執行極度依賴複雜硬體(如保留站 Reservation Station、發射佇列 Issue Queue、暫存器重命名單元、重排序緩衝區 ROB 等)來進行動態排程。
  3. 精確例外(Precise Exceptions)與重排序緩衝區(Reorder Buffer, ROB):
    • 要求例外發生時的架構狀態(Architectural State)必須與依序執行(In-order Execution)的結果完全一致。
    • 處理器採取「亂序執行(Execute Out-of-Order)、依序提交(Commit / Retire In-Order)」策略。
  4. 投機執行(Speculative Execution)與誤預測恢復:
    • 分支預測錯誤時,必須撤銷(Flush/Cancel)所有沿著錯誤跳轉路徑抓取且**尚未提交(Uncommitted)**的投機指令。已提交(Committed)的指令已正式修改架構狀態,不可撤銷。

解題方法

本題為微架構觀念判斷題。解題切入點為檢視現代亂序超純量處理器的標準流水線運作流程:
Fetch→Decode / Rename→Issue / Reservation Station→Execute (OoO)→Writeback→Commit / Retire (In-Order via ROB)\text{Fetch} \rightarrow \text{Decode / Rename} \rightarrow \text{Issue / Reservation Station} \rightarrow \text{Execute (OoO)} \rightarrow \text{Writeback} \rightarrow \text{Commit / Retire (In-Order via ROB)}

根據各階段的硬體責任、資料流約束及狀態更新機制,逐一分析五個選項的敘述是否符合教科書標準定義。


選項分析

  • (a) 正確
    在傳統流水線中,分支指令需等到 Decode 或 Execute 階段才能確定跳轉方向與目標位址,這會造成控制冒險並產生管道停頓(Pipeline Stalls)。
🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 5 題10 分

  1. (10pts, no partial point, no penalty) A processor runs at a clock frequency of 2 GHz. A particular program executes 1 billion instructions (10^9) with the following instruction mix and clock-cycle counts per instruction type:
  • 20% R-type instructions, each requiring 4 cycles
  • 50% I-type instructions, each requiring 5 cycles
  • 30% Memory (load/store) instructions, each requiring 8 cycles
    Which of the following statements correctly gives both the effective CPI and the total CPU execution time for this program?
    (a) Effective CPI = 5.2 and CPU time = 2.60 seconds.
    (b) Effective CPI = 5.7 and CPU time = 2.85 seconds
    (c) Effective CPI = 5.7 and CPU time = 1.70 seconds
    (d) Effective CPI = 6.2 and CPU time = 3.10 seconds
    (e) Effective CPI = 5.0 and CPU time = 2.50 seconds

登入後即可作答並保存紀錄。

這一題的完整詳解

核心觀念

本題考查《計算機組織與結構》中 CPU 效能方程式(CPU Performance Equation) 的核心概念,主要包含以下兩個基礎公式:

  1. 加權平均有效 CPI(Effective Cycles Per Instruction):
    當程式包含多種不同指令類型且各自所需的時脈週期數(CPI)不同時,整體程式的有效 CPI 為各指令類型出現比例與其對應 CPI 的加權平均值:
    CPIeffective=∑i=1n(fi×CPIi)\text{CPI}_{\text{effective}} = \sum_{i=1}^{n} (f_i \times \text{CPI}_i)
    其中 fif_i 為第 ii 種指令類型的出現比例(Instruction Mix),CPIi\text{CPI}_i 為該指令類型的時脈週期數。

  2. CPU 執行時間(CPU Execution Time):
    程式在 CPU 上的總執行時間取決於指令總數(Instruction Count)、有效 CPI 與時脈頻率(Clock Frequency):
    CPU Execution Time=Instruction Count×CPIeffectiveClock Frequency(f)\text{CPU Execution Time} = \frac{\text{Instruction Count} \times \text{CPI}_{\text{effective}}}{\text{Clock Frequency} (f)}


解題方法

根據題目給定條件:

  • 時脈頻率 f=2 GHz=2×109 Hzf = 2 \text{ GHz} = 2 \times 10^9 \text{ Hz}
  • 指令總數 Instruction Count=1 billion=109\text{Instruction Count} = 1 \text{ billion} = 10^9
  • 各類別指令比例與其對應 CPI:
    • R-type 指令:比例 fR=20%=0.20f_{\text{R}} = 20\% = 0.20,CPIR=4\text{CPI}_{\text{R}} = 4
    • I-type 指令:比例 fI=50%=0.50f_{\text{I}} = 50\% = 0.50,CPII=5\text{CPI}_{\text{I}} = 5
    • Memory (load/store) 指令:比例 fMem=30%=0.30f_{\text{Mem}} = 30\% = 0.30,CPIMem=8\text{CPI}_{\text{Mem}} = 8

步驟一:計算有效 CPI (CPIeffective\text{CPI}_{\text{effective}})

將各類別指令的比例及其對應的 CPI 代入加權平均公式:

🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 6 題10 分

  1. (10pts, no partial point, no penalty) In a modern computer system with multiple levels of caches (L1, L2, and possibly L3), which of the following statements about cache design and performance is/are FALSE?
    (a) Increasing the cache block size (line size) beyond a certain point can reduce the miss rate, but it may also increase the miss penalty due to more data transfer on each miss.
    (b) Write-back caches always require a write to main memory on every store instruction, ensuring data consistency at the expense of performance.
    (c) The inclusion property states that everything in the L1 cache must also be present in the L2 cache, ensuring consistent copies of data across levels.
    (d) Associative caches (with more ways) generally reduce the conflict miss rate but can have higher access latency and more complex hardware.
    (e) A fully associative cache of the same size as a direct-mapped cache will always have a lower miss rate, but its hardware complexity and access time might be higher.

登入後即可作答並保存紀錄。

這一題的完整詳解

核心觀念

本題考驗現代電腦系統中多層次快取(Multi-level Cache)設計與效能分析的核心知識,包含以下觀念與定量/定性原則:

  1. 快取失誤分類(3Cs Model)與效能因子:

    • Compulsory Miss(強制失誤):首次存取資料塊時產生的失誤,無法藉由改變快取結構避免。
    • Capacity Miss(容量失誤):快取總容量不足以容納程式所需的所有資料塊而產生的失誤。
    • Conflict Miss(衝突失誤):多個資料塊映射到相同 Set 而提早被替換產生的失誤。
  2. Block Size(區塊大小)對效能的雙重效應:

    • 在快取總容量固定下,適度增加 Block Size 可發揮**空間局部性(Spatial Locality)**以降低 Miss Rate。
    • 若 Block Size 超過一定臨界點,區塊總數量(Number of blocks=Cache SizeBlock Size\text{Number of blocks} = \frac{\text{Cache Size}}{\text{Block Size}})過少,會導致區塊競爭加劇,Miss Rate 反而上升;同時每次 Miss 需傳輸更多資料,使 Miss Penalty 變大。
  3. 快取寫入策略(Write Policy):

    • Write-through(直寫):每次 Store 指令皆同時寫入 Cache 與 Main Memory。
    • Write-back(寫回):Store 指令僅更新 Cache 並設置 Dirty bit,直到該區塊被替換(Evict)時才寫回 Main Memory。
  4. 包含屬性(Inclusion Property):

    • Inclusive Cache 規定 L1 快取內的所有資料必須同時存在於 L2 快取中(L1⊂L2L1 \subset L2),有利於多核心匯流排監聽(Bus Snooping)維持 Consistency。
  5. 相聯度(Associativity)與硬體代價:

    • 提高相聯度(Direct-mapped →\rightarrow Set-associative →\rightarrow Fully associative)可大幅減少 Conflict Miss。
    • 但需要更多比較器(Comparators)與多工器(Multiplexers),會增加硬體複雜度與存取延遲(Hit Time / Access Latency)。

解題方法

本題為複選題型(選出所有**錯誤(FALSE)**的敘述)。解題切入點如下:

  1. 比對機制定義:核對 Write-back 與 Write-through 的差異,以及 Inclusive Cache 的定義。
  2. 分析極限與趨勢:檢視 Block Size 過大時對 Miss Rate 與 Miss Penalty 的影響。
  3. 檢驗絕對性敘述:針對帶有「always」等過於絕對特性的敘述,尋找反例(如強制失誤主導或病態存取模式)加以否決。

選項分析

  • (a) 錯誤(FALSE)

    • 選項原文:Increasing the cache block size (line size) beyond a certain point can reduce the miss rate, but it may also increase the miss penalty due to more data transfer on each miss.
    • 詳盡剖析:適度增加 Block Size 能利用空間局部性降低失誤率。但當 Block Size 超過一定臨界點(beyond a certain point) 時,快取內的區塊總數過少,導致衝突與容量失誤大幅增加,Miss Rate 反而會上升(increase),而非繼續降低(reduce)。此外,每次失誤傳輸資料量變大,確實會增加 Miss Penalty。選項敘述「超過臨界點還能降低 Miss Rate」與快取效能理論不符。
  • (b) 錯誤(FALSE)

    • 選項原文:Write-back caches always require a write to main memory on every store instruction, ensuring data consistency at the expense of performance.
    • 詳盡剖析:每次執行 Store 指令都必須寫回 Main Memory 的機制為 Write-through(直寫);
🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 7 題10 分

  1. (10pts, no partial point, no penalty) A modern processor may implement multi-core designs (e.g., 4 cores), and each core may support hardware-based multithreading (e.g., 2 hardware threads per core). Which of the following statements about multicore and hardware multithreading is/are FALSE?
    (a) Hardware-based multithreading (like Simultaneous Multithreading, SMT) allows each core to issue instructions from multiple threads in the same clock cycle, when resources (functional units) are available.
    (b) When the CPU supports 2 hardware threads per core on a 4-core processor, the operating system sees a total of 4 logical CPUs in the system.
    (c) Multicore processors place multiple independent processing units (cores) on a single chip, each capable of running its own thread of execution.
    (d) A multicore design can achieve parallel execution even if hardware multithreading is disabled or not implemented, because each core can run an independent process or thread.
    (e) Hardware-based multithreading helps hide pipeline stalls (e.g., due to memory latency) by quickly switching to another hardware thread, utilizing idle execution resources.

登入後即可作答並保存紀錄。

這一題的完整詳解

核心觀念

本題考驗現代處理器微架構中**多核心(Multicore)與硬體多執行緒(Hardware Multithreading)**的定義、運作機制及其與作業系統(OS)間的對映關係。

關鍵定義與公式:

  1. 多核心(Multicore Processor):在單一矽晶片(Chip)上整合多個獨立的實體處理核心(Physical Cores)。每個核心皆擁有完整的執行管線(Pipeline)與運算的算術邏輯單元(ALU),可同時獨立平行執行不同的行程或執行緒。
  2. 硬體多執行緒(Hardware Multithreading):在單一實體核心內部重複配備暫存器檔(Register File)、程式計數器(PC)等架構狀態(Architectural State)。當前執行緒因記憶體存取或管線冒險而停頓(Stall)時,能快速切換並利用空閒的執行資源。
    • 同時多執行緒(Simultaneous Multithreading, SMT):如 Intel 的 Hyper-Threading 技術,允許在同一時脈週期內,發射來自多個不同執行緒的指令至獨立的功能單元(Functional Units)。
  3. 邏輯 CPU 數量公式:
    邏輯 CPU 總數 (Logical CPUs)=實體核心數 (Physical Cores)×每核心硬體執行緒數 (Hardware Threads per Core)\text{邏輯 CPU 總數 (Logical CPUs)} = \text{實體核心數 (Physical Cores)} \times \text{每核心硬體執行緒數 (Hardware Threads per Core)}
    作業系統視每一個硬體執行緒為一個獨立的邏輯 CPU(Logical Processor),並據此實施多工作業與執行緒排程。

解題方法

題目要求選出關於多核心與硬體多執行緒敘述**錯誤(FALSE)**的選項。切入步驟如下:

  1. 計算系統的邏輯 CPU 數量:將題目給定之 4 個實體核心與每核心 2 個硬體執行緒代入公式:
    邏輯 CPU 總數=4×2=8\text{邏輯 CPU 總數} = 4 \times 2 = 8
    作業系統應識別到 8 個邏輯 CPU,與選項 (b) 敘述的 4 個不符。
  2. 驗證其餘架構概念:依據電腦組織學定義,逐一核對 (a)、(c)、(d)、(e) 之敘述,確認其皆符合作業系統與微架構的實務運作機制。
🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 8 題10 分

  1. (10pts, no partial point) Assume a simple 5-stage RISC-V pipeline (IF → ID → EX → MEM → WB) with:
  • Forwarding from MEM to EX (to reduce some data hazards).
  • No branch prediction, and branches are resolved in EX (meaning the pipeline does not know if the branch is taken or not until the EX stage of the branch instruction).
    Consider the following snippet of RISC-V assembly:
    L1:
    lw x5, 0(x10)
    add x6, x5, x7
    sub x8, x6, x5
    beq x8, x0, L1
    or x9, x8, x0
    add x9, x0, x0
    How many stall cycles (pipeline bubbles) are introduced in total by these instructions due to hazards, under the given assumptions?

登入後即可作答並保存紀錄。

這一題的完整詳解

解題步驟

  1. 資料冒險 (Data hazards)
    • lw x5,0(x10) → add x6,x5,x7 為 load‑use。即使有 MEM→EX forwarding,資料在 lw 的 MEM 階段才取得,add 的 EX 需要在同一週期使用,仍必須插入 1 個泡沫。
    • add x6,x5,x7 → sub x8,x6,x5
      • 來源 x6 可於 add 進入 MEM 時轉送至 sub 的 EX,無需停等。
      • 來源 x5 已在 lw 完成寫回,亦無停等。
    • sub x8,x6,x5 → beq x8,x0,L1
      • x8 於 sub 的 MEM 階段可轉送至 beq 的 EX,無停等。
🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 9 題10 分

(10pts, no partial point) A processor has:

  • A 32-bit virtual address space with a two-level page table.
  • A page size of 4 KB (4096 bytes).
  • A single TLB (Translation Lookaside Buffer) with:
    • TLB hit rate = 98% (so TLB miss rate = 2%).
    • Access time for a TLB hit is already included (hidden) in the base pipeline.
    • Miss penalty: On a TLB miss, the hardware must walk the two-level page table in memory. Each level of the page table is stored in main memory, which takes 50 cycles per memory access.

Other system parameters:

  • The base CPI (assuming no TLB misses) is 1.0.
  • 40% of the instructions perform a data access (load/store) in addition to the instruction fetch.
  • Each instruction fetch needs a TLB translation, and each load/store also needs a TLB translation.
  • Ignore any other stall sources (cache misses, etc.). Focus only on the extra stalls from TLB misses.

What is the overall (effective) CPI once TLB misses and their penalties are included?

Overall CPI = ____.

登入後即可作答並保存紀錄。

這一題的完整詳解

核心觀念

有效 CPI 等於基礎 CPI 加上每條指令平均承受的 TLB 額外停滯週期:

有效 CPI=基礎 CPI+每指令平均 TLB miss 次數×每次 miss 的停滯週期\text{有效 CPI} = \text{基礎 CPI} + \text{每指令平均 TLB miss 次數} \times \text{每次 miss 的停滯週期}

每次 TLB miss 都需要依序查詢兩層頁表;頁表每層各需一次記憶體存取。

解題方法

依原卷第 4 頁第 9 題,處理器採用兩層頁表、TLB 命中率為 98%98\%,每次記憶體存取需 5050 個週期;每條指令都要進行一次指令擷取,另有 40%40\% 的指令會進行一次資料存取。基礎 CPI 為 1.01.0,且命中 TLB 的時間已包含在基礎管線中。

先計算一次 TLB miss 的停滯週期:

2 次頁表存取×50 週期/次=100 週期2 \text{ 次頁表存取} \times 50 \text{ 週期/次} = 100 \text{ 週期}

每條指令平均需要的位址轉譯次數為:

🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 10 題10 分

(10pts) A system uses 32-bit addresses and has an 8 KB (8192 bytes) 2-way set-associative cache. The block (line) size is 16 bytes. Answer the following question

  • number of sets = ____.
  • Number of offset bits = ____.
  • The number of index bits = ____.
  • The number of tag bits = ____.

登入後即可作答並保存紀錄。

這一題的完整詳解
載入中…

其他考古題