112 年 國立清華大學計量財務金融研究所財務金融組《統計學》
第 1 題20 分
- The probability that a person has a certain disease is 0.05. Medical diagnostic tests are available to determine whether the person actually has the disease. If the disease is actually present, the probability that the medical diagnostic test will give a positive result (indicating that the disease is present) is 0.85. If the disease is not actually present, the probability of a positive test result (indicating that the disease is present) is 0.03. Suppose that the medical diagnostic test has given a positive result (indicating that the disease is present). What is the probability that the disease is actually present? (20%)
登入後即可作答並保存紀錄。
此題考驗貝氏定理(Bayes' Theorem)的應用,主要在於計算後驗機率。
設 D 為該人患有此疾病的事件,D' 為該人未患有此疾病的事件。
設 T 為醫學診斷測試結果為陽性的事件。
根據題目敘述,我們有以下機率:
P(D) = 0.05 (患有此疾病的先驗機率)
P(D') = 1 - P(D) = 1 - 0.05 = 0.95 (未患有此疾病的先驗機率)
P(T|D) = 0.85 (若疾病存在,測試結果為陽性的機率,即靈敏度)
P(T|D') = 0.03 (若疾病不存在,測試結果為陽性的機率,即假陽性率)
題目要求計算的是,在測試結果為陽性的條件下,該人實際上患有此疾病的機率,即 P(D|T)。
根據貝氏定理:
第 2 題20 分
- (20%) A coin-operated soft-drink machine is designed to discharge at least 8 ounces of beverage per cup, with a standard deviation of 0.16 ounce. If you select a random sample of 36 cups and you are willing to have an risk of committing a Type I error, compute the probability of a Type II error if the population mean amount dispensed is actually 7.9 ounces per cup. (20%)
登入後即可作答並保存紀錄。
此題考驗假設檢定中第一類錯誤與第二類錯誤的機率計算,特別是第二類錯誤的機率 。
首先,我們需要設定虛無假設 和對立假設 。
題目提到機器設計為「至少 8 盎司」(at least 8 ounces),這表示我們檢定的目標是平均量是否達到 8 盎司。
所以,虛無假設應設定為平均量等於或大於 8 盎司,而對立假設為平均量小於 8 盎司。
(或寫成 ,因為檢定時通常以等號為界)
這是一個左尾檢定(left-tailed test)。
已知資訊:
母體標準差 盎司 (題目未明確說明是母體標準差還是樣本標準差,但通常在這種情境下,若未提及樣本標準差,且樣本數較大,常 Assume 為母體標準差,或者題目可能意指母體標準差已知。若為樣本標準差,則應使用 t-test。但由於題目給了 且要求計算 Type II error,通常涉及 Z-score,暗示 已知或可用樣本標準差近似。)
樣本大小
第一類錯誤的機率
在實際情況下,母體平均值 盎司。
第一步:確定拒絕域 (Rejection Region)
我們進行左尾檢定,顯著水準為 。
由於樣本數 ,我們可以根據中央極限定理使用 Z-檢定。
拒絕域為 。
查標準常態分配表,當 時,。
因此,若檢定統計量 ,則拒絕 。
檢定統計量公式為:
其中 是虛無假設下的母體平均值,在此為 8。
拒絕域的條件 值:
第 3 題20 分
- (20%) The following table shows the joint distribution of two discrete random variables, X and Y.
| Y=-2 | Y=0 | Y=2 | |
|---|---|---|---|
| X=-1 | 0 | 0.1 | 0.2 |
| X=0 | 0 | 0.2 | 0.1 |
| X=3 | 0.3 | 0.1 | 0 |
Compute the correlation coefficient between and . (20%)
登入後即可作答並保存紀錄。
此題考驗計算兩個離散隨機變數相關係數的能力,並涉及變數的線性轉換。
相關係數 的定義為:
其中 是 X 和 Y 的共變異數, 和 分別是 X 和 Y 的標準差。
題目要求計算的是 和 這兩個新變數之間的相關係數。
設 且 。
我們需要計算 。
一個重要的性質是,對變數進行線性轉換,並不會改變它們之間的相關性。
也就是說,對於任意常數 ,其中 ,隨機變數 與 的相關係數為:
如果 和 同號 (都為正或都為負),則 ,所以 。
如果 和 異號,則 ,所以 。
在本題中, (), ()。因為 和 都是正數,所以:
因此,我們只需要計算 X 和 Y 的相關係數 。
要計算 ,我們需要計算 、 和 。
1. 計算邊際機率分佈 (Marginal Probability Distributions)
X 的邊際機率分佈:
(檢查總和:0.3 + 0.3 + 0.4 = 1.0)
Y 的邊際機率分佈:
(檢查總和:0.3 + 0.4 + 0.3 = 1.0)
2. 計算期望值 E(X) 和 E(Y)
第 4 題20 分
- The following information is available for two samples selected from independent populations: (At the 0.05 level of significance)
Sample 1: n1 = 4, 3, 5, 6, 11, 7, 8.2, 9.1, 12, 14, 14.5
Sample 2: n2 = 9, 10.3, 12.5, 16.7, 18, 6.3, 11.5, 2, 12.4, 19.3, 21.5
Replace the observed values with the corresponding ranks (where smallest value=1; largest value= n = n1 + n2) in the combined samples. At the 0.05 level of significance, is there evidence of a difference in the median waiting time between the two samples? (Since both of the sample sizes are greater than 10, use the large-sample Z approximation formula) (20%)
登入後即可作答並保存紀錄。
核心觀念
- 本題檢驗的是兩個獨立樣本的中位數是否相同,屬於 Wilcoxon (Mann‑Whitney) 排名和檢定。
- 當兩樣本皆 > 10 時,可用 大樣本 Z 近似:
其中
解題方法
- 合併樣本並排序(最小值 rank = 1,最大值 rank = )。
- 計算樣本 1 的秩和 ,再求其 Mann‑Whitney 。
- 用 計算 Z 值,與標準正態臨界值 比較(雙側檢定),決定是否拒絕原假設 兩樣本中位數相等。
以下列出完整步驟與計算。
1. 排名 (Rank)
| 級別 | 原始值 | 所屬樣本 | Rank |
|---|---|---|---|
| 1 | 2 | Sample 2 | 1 |
| 2 | 3 | Sample 1 | 2 |
| 3 | 4 | Sample 1 | 3 |
| 4 | 5 | Sample 1 | 4 |
| 5 | 6 | Sample 1 | 5 |
| 6 | 6.3 | Sample 2 | 6 |
| 7 | 7 | Sample 1 | 7 |
| 8 | 8.2 | Sample 1 | 8 |
| 9 | 9 | Sample 2 | 9 |
| 10 | 9.1 | Sample 1 | 10 |
| 11 | 10.3 | Sample 2 | 11 |
| 12 | 11 | Sample 1 | 12 |
| 13 | 11.5 | Sample 2 | 13 |
| 14 | 12 | Sample 1 | 14 |
| 15 | 12.4 | Sample 2 | 15 |
| 16 | 12.5 | Sample 2 | 16 |
| 17 | 14 | Sample 1 | 17 |
| 18 | 14.5 | Sample 1 | 18 |
| 19 | 16.7 | Sample 2 | 19 |
| 20 | 18 | Sample 2 | 20 |
| 21 | 19.3 | Sample 2 | 21 |
第 5 題20 分
- Suppose two professors wanted to study how students from their two universities (A and B) compared in their capabilities of using Excel spreadsheets in undergraduate information systems courses. A comparison of the student demographics was also performed. The following table contains information regarding the ages of the students:
| School | Sample size | Mean | Standard deviation |
|---|---|---|---|
| A | 9 | 23.23 | 2.23 |
| B | 13 | 21.12 | 1.54 |
a. Compute the whole sample variation. (5%)
b. To test for a difference in the mean age of students, is it most appropriate to use the pooled-variance t test or the separate-variance t test? Use the most appropriate test to determine, at the 0.05 level of significance, whether there is evidence of a difference in the mean years of spreadsheet usage of students at A school and at B school.(15%)
登入後即可作答並保存紀錄。
此題包含兩個部分:(a) 計算全體樣本的變異數,(b) 進行獨立樣本 t 檢定,並判斷應使用合併變異數 t 檢定或獨立變異數 t 檢定。
a. 計算全體樣本的變異數 (Whole Sample Variation)
全體樣本的變異數 (pooled variance) 是指將兩個樣本合併視為一個大樣本後計算出的變異數。通常在進行合併變異數 t 檢定時會用到。
我們有兩個樣本的資訊:
樣本 A: , ,
樣本 B: , ,
首先,計算兩個樣本的平方和 (sum of squares):
合併後的樣本數 。
合併後的總平方和 。
全體樣本的變異數 (pooled variance, ) 是總平方和除以合併後的自由度:
自由度 。
另一種計算全體樣本變異數的觀點:
題目問的是 "whole sample variation",這通常是指合併兩個樣本後的變異數。
合併變異數 的公式為:
解釋為何有兩種結果:
第一種計算方式 (3.2496) 是將兩個樣本的變異數加總後,除以總自由度 。這更像是描述整個數據集的總離散度。
第二種計算方式 (3.41212) 是計算合併變異數 的標準公式,用於合併變異數 t 檢定。
題目要求 "whole sample variation",這比較模糊。在統計學中,"variation" 通常指平方和 (SS) 或變異數 (variance)。如果指變異數,則 是最常見的「合併樣本變異數」。如果指「全體樣本的變異數」,那麼 也是一種解釋。
然而,在進行 t 檢定時,我們使用的是 。
我將採用這個公式的結果作為答案。
a. 計算結果:
全體樣本變異數 (合併變異數) 。
b. 判斷應使用哪種 t 檢定,並進行檢定
判斷依據:
當我們比較兩個獨立樣本的平均數時,有兩種主要的 t 檢定:
- 合併變異數 t 檢定 (Pooled-variance t test): 假設兩個母體的變異數相等 ()。如果這個假設成立,合併變異數 的估計會更有效率。
- 獨立變異數 t 檢定 (Separate-variance t test, Welch's t-test): 不假設兩個母體變異數相等。這個檢定更為穩健,因為它不需要變異數相等的假設,並且在變異數不相等時,其自由度會被調整。
如何判斷?
通常,我們會先對變異數是否相等進行檢定 (例如 F 檢定)。但題目要求我們「最適合」使用哪一種。
觀察樣本變異數的估計值:
樣本 A 的變異數 約為樣本 B 的變異數 的兩倍。
雖然 比 大,但由於樣本數不同 (),我們需要更嚴謹的檢定來判斷母體變異數是否相等。
然而,在沒有進行變異數相等檢定的情況下,如果兩者變異數差異較大,或者樣本數差異較大,獨立變異數 t 檢定 (Welch's t-test) 通常被認為是更安全的選擇,因為它對變異數不相等的假設不敏感。
但題目問「最適合」,這可能意味著要考慮兩種檢定的優勢。如果變異數相等,合併變異數 t 檢定有較高的檢定力。
這裡的關鍵在於題目給的數據:
, 。
, 。
。
這個比例並不算極大。