108 年 國立臺灣大學土木系碩士班交通組《統計學(C)》
第 1 題10 分
- (10 points) Show that the mean of a hypergeometric random variable X, the number of successes in a random sample of size selected from items of which are labeled success and labeled failure, is .
登入後即可作答並保存紀錄。
此題考驗對超幾何分布期望值的理解。超幾何分布描述了從一個有限母體中進行不放回抽樣時,抽中特定屬性(成功)的次數的機率分布。
證明過程如下:
令 為超幾何隨機變數,代表從 個物品中抽取 個樣本,其中有 個是成功的, 個是失敗的。
超幾何分布的機率質量函數 (PMF) 為:
根據期望值的定義,我們有:
我們知道 。將此性質代入上式:
將常數 和 提出求和符號:
現在,我們需要化簡 。我們知道 。
第 2 題10 分
- (10 points) The amount of lateral expansion (mils) was determined for a sample of pulsed-power gas metal arc welds used in LNG ship containment tanks. The resulting sample standard deviation was mils. Assuming normality, derive a 95% confidence intervals for .
登入後即可作答並保存紀錄。
此題考驗對卡方 (Chi-squared, ) 分布的理解,以及如何建構母體變異數 的信賴區間。
核心觀念:
當母體服從常態分布且其變異數未知時,樣本變異數 與母體變異數 的關係,可以透過卡方分布來描述。具體而言,統計量 服從自由度為 的卡方分布。
題目資訊:
- 樣本大小
- 樣本標準差
- 樣本變異數
- 信賴水準為 95%
推導步驟:
-
確定統計量:
根據假設,母體服從常態分布。統計量 服從自由度為 的卡方分布。
在此題中,自由度 。
因此,。 -
建立信賴區間的機率不等式:
對於 95% 的信賴水準,我們需要在卡方分布的兩側各留出 的機率。
令 和 分別代表自由度為 的卡方分布的右側 和左側 的臨界值(或左側 的臨界值)。
則有:
將 和 代入:
第 3 題10 分
- (10 points) Consider the following frequency table of observations on the random variable X.
Value | 0 | 1 | 2 | 3 | 4 | 5
------- |---|---|---|---|---|---
Observed Frequency | 8 | 25 | 23 | 21 | 16 | 7
(a) Based on these 100 observations, is a Poisson distribution with a mean of 2.4 an appropriate model? Perform a goodness-of-fit test with .
(b) Perform the test using P-value approach.
登入後即可作答並保存紀錄。
此題考驗對卡方 (Chi-squared, ) 適合度檢定 (Goodness-of-Fit Test) 的應用,以及 P-value 的計算與解釋。
核心觀念:
適合度檢定用於判斷觀測到的數據是否符合某一個理論上的機率分布。卡方適合度檢定是其中一種常用方法,它比較觀測頻率 (Observed Frequency, ) 與期望頻率 (Expected Frequency, ) 之間的差異。
題目資訊:
- 觀測數據:給定的頻率表,總觀測數 。
- 檢定假設:母體服從平均數為 的 Poisson 分布。
- 顯著水準 。
步驟一:計算樣本平均數(以確認 Poisson 分布的參數)
首先,我們需要根據觀測數據計算樣本的平均數,因為 Poisson 分布的參數 通常等於母體平均數。
樣本平均數
其中 是對應於數值 的觀測頻率。
題目要求檢定的是一個 指定參數 的 Poisson 分布,即 。雖然樣本平均數 與 相近,但我們仍需使用題目指定的 來計算期望頻率。
步驟二:計算理論 Poisson 分布的機率與期望頻率
假設 。
Poisson 分布的機率質量函數 (PMF) 為 。
我們需要計算 對於 。
合併類別(重要):
在進行卡方適合度檢定時,每個類別的期望頻率 至少應該為 5。
讓我們計算期望頻率:
所有計算出的期望頻率都大於 5,因此不需要合併類別。
步驟三:計算卡方檢定統計量
卡方檢定統計量為:
| Value () | Observed () | Expected () |
|---|
第 4 題20 分
- (20 points) An article describes a new equivalent plate analysis method formulation that is capable of modeling aircraft structures such as cranked wing boxes and that produces results similar to the more computationally intensive finite element analysis method. Natural vibration frequencies for the cranked wing box structure are calculated using both methods, and results for the first seven natural frequencies are shown here.
| Car | Finite element, cycle/s | Equivalent Plate, cycle/s | Car | Finite element, cycle/s | Equivalent Plate, cycle/s |
|---|---|---|---|---|---|
| 1 | 14.58 | 14.76 | 5 | 174.73 | 181.22 |
| 2 | 48.52 | 49.10 | 6 | 212.72 | 220.14 |
| 3 | 97.22 | 99.99 | 7 | 277.38 | 294.80 |
| 4 | 113.99 | 117.53 |
(a) Do the data suggest that the two methods provide the same mean value for natural vibration frequency? Using the P-value approach.
(b) Find a 95% confidence interval on the mean difference between the two methods and use it to answer the question in part (a).
登入後即可作答並保存紀錄。
此題考驗對配對樣本 t 檢定 (Paired t-test) 和配對樣本信賴區間的應用。題目要求比較兩種量測方法(Finite element 和 Equivalent Plate)的平均值是否有顯著差異。
核心觀念:
當我們比較兩個相關樣本(例如,同一物體在不同條件下的量測,或同一對象的兩種不同量測方法)的平均值時,我們使用配對樣本 t 檢定。配對樣本 t 檢定的基本思想是計算每一對觀測值之間的差值,然後對這些差值進行單一樣本 t 檢定。
步驟一:定義差值變數
令 為第 次觀測時,兩種方法測量結果的差值。
在此,我們定義 。
這樣定義可以讓我們檢定「兩種方法是否提供相同的平均值」,即檢定差值的平均數是否為零。
計算差值:
步驟二:計算差值樣本的平均數 和樣本標準差
樣本大小 (因為有 7 對觀測值)。
計算樣本平均數 :
計算樣本標準差 :
首先,計算樣本變異數 。
:
:
第 5 題10 分
- (10 points) Let X be a random variable with distribution function
Plot the graph of F(x) and then evaluate E(x).
登入後即可作答並保存紀錄。
此題考驗對累積分布函數 (CDF) 的圖形繪製以及計算隨機變數期望值 (E(X)) 的能力。
步驟一:繪製累積分布函數 F(x) 的圖形
CDF 的定義是 。
函數 有兩個部分:
- 當 時,。這表示在 x 軸的負半軸,圖形是水平的,值為 0。
- 當 時,。
- 當 時,。
- 隨著 趨近於無限大 (), 趨近於 0,所以 趨近於 。
- 是一個指數遞增的曲線,從 0 開始,漸近於 1。
圖形描述:
- 在 的區域,圖形是 x 軸本身。
- 在 處,圖形從 (0, 0) 開始(包含點 (0,0))。
- 對於 ,圖形是一個平滑的曲線,向上遞增,並漸近於水平線 。
- 這是一個標準的指數分布的 CDF 圖形。
步驟二:計算期望值 E(X)
首先,我們需要找到機率密度函數 (PDF) 。PDF 是 CDF 的導數。
- 對於 ,,所以 。
- 對於 ,,所以 。
因此,PDF 為:
這是一個參數為 的指數分布。
期望值的定義是 。
第 6 題10 分
- (10 points) There are two classes of Probability and Statistics course. The problems of the final exam are identical for the two classes. Let and denote the population means of scores of class A and class B, respectively, and and denote the population variances of scores of class A and class B, respectively. A random sample of 13 students are taken from class A and 11 students are taken from class B. The sample average and standard deviation of scores of each class are class A: and class B: . Test whether the population means of scores of the two classes are equal. Make assumptions if necessary.
登入後即可作答並保存紀錄。
此題考驗對兩獨立樣本平均數比較檢定 (Two-sample t-test) 的應用。題目要求檢定兩個班級的平均分數是否相等。
核心觀念:
當我們要比較兩個獨立樣本的平均值時,我們使用兩獨立樣本 t 檢定。此檢定有兩種常見形式,取決於我們是否假設兩個母體的變異數相等:
- 假設 (Pooled variance t-test): 當我們有理由相信兩個母體的變異數相等時使用。
- 不假設 (Welch's t-test): 當我們不能確定母體變異數相等時使用。
題目要求「Make assumptions if necessary」,這表示我們可以自行選擇是否假設變異數相等。通常,如果樣本大小差異不大且樣本標準差相近,可以考慮池化變異數;否則,Welch's t-test 更為穩健。在此,我們將採用 Welch's t-test,因為它不需要假設變異數相等。
題目資訊:
- 班級 A:
- 樣本大小
- 樣本平均數
- 樣本標準差
- 班級 B:
- 樣本大小
- 樣本平均數
- 樣本標準差
步驟一:設定假設
- 虛無假設 : 兩個班級的母體平均分數相等,即 或 。
- 對立假設 : 兩個班級的母體平均分數不相等,即 或 。
這是一個雙尾檢定。
步驟二:選擇檢定方法與計算檢定統計量
我們採用 Welch's t-test,它不假設母體變異數相等。
檢定統計量為:
其中 根據虛無假設。
計算樣本變異數:
代入數值:
步驟三:計算自由度 (Welch-Satterthwaite 方程式)
Welch's t-test 的自由度計算較複雜,使用 Welch-Satterthwaite 方程式:
第 7 題15 分
- (15 points) Consider the experiment of tossing two dices. Let X denote the sum of the numbers that occur. Let Y denote the absolute difference of the numbers that occur.
(a) What is the probability function of Y?
(b) What is the cumulative distribution function of Y?
(c) Find the mean and variance of X.
登入後即可作答並保存紀錄。
此題考驗對離散隨機變數機率分布的建立、累積分布函數的求解,以及期望值與變異數的計算。
實驗設定: 投擲兩個公正的骰子。
樣本空間總共有 種可能結果。
令第一個骰子的點數為 ,第二個骰子的點數為 。
定義隨機變數:
- (兩個骰子點數的總和)
- (兩個骰子點數的絕對差值)
(a) 求 Y 的機率函數 (Probability Function, PMF)
的可能值為何?
的可能值為:
- 0 (例如:(1,1), (2,2), ..., (6,6))
- 1 (例如:(1,2), (2,1), (2,3), (3,2), ...)
- 2 (例如:(1,3), (3,1), (2,4), (4,2), ...)
- 3 (例如:(1,4), (4,1), (2,5), (5,2), ...)
- 4 (例如:(1,5), (5,1), (2,6), (6,2))
- 5 (例如:(1,6), (6,1))
所以 的可能值為 。
我們需要計算 對於 。
-
: 。
有利結果:(1,1), (2,2), (3,3), (4,4), (5,5), (6,6)。共有 6 種。
。 -
: 。
有利結果:(1,2), (2,1), (2,3), (3,2), (3,4), (4,3), (4,5), (5,4), (5,6), (6,5)。共有 10 種。
。 -
: 。
有利結果:(1,3), (3,1), (2,4), (4,2), (3,5), (5,3), (4,6), (6,4)。共有 8 種。
。 -
: 。
有利結果:(1,4), (4,1), (2,5), (5,2), (3,6), (6,3)。共有 6 種。
。 -
: 。
有利結果:(1,5), (5,1), (2,6), (6,2)。共有 4 種。
。 -
: 。
有利結果:(1,6), (6,1)。共有 2 種。
。
檢查總機率:
。
Y 的機率函數 (PMF):
簡化分數:
第 8 題15 分
- (15 points) The data of the thrust of a jet-turbine engine (y), =primary speed of rotation, =secondary speed of rotation, and =fuel flow rate are collected. It is known that all three independent variables (x) have positive effect on the dependent variable (y). The candidate regression model for the data is . Answer the following questions based on the Excel report shown below.
Regression Statistics
Multiple R | 0.996618
R Square | 0.993248
Adjusted R Square | 0.992685
Standard Error | 43.16167
Observations | 40
ANOVA
| df | SS | MS | F | Significance F | |
|---|---|---|---|---|---|
| Regression | 3 | 9864971 | 3288324 | 1765.135 | 4.15E-39 |
| Residual | 36 | 67065.48 | 1862.93 | ||
| Total | 39 | 9932036 |
| Coefficients | Std Error | t Stat | P-values | Lower 95% | Upper 95% |
|---|---|---|---|---|---|
| Intercept | 5763.816 | 1324.209 | 4.352647 | 0.000106 | 3078.195 |
| x1 | 2.162198 | 0.162787 | 13.28236 | 1.92E-15 | 1.832051 |
| x2 | 0.068896 | 0.058009 | 1.187672 | 0.24274 | -0.04875 |
| x3 | -0.24037 | 0.064567 | -3.72279 | 0.000671 | -0.37132 |
(a) What is the fitted linear regression model?
(b) What is the coefficient of determination of this model?
(c) Does the model coefficients consistent with the prior knowledge that all three independent variables (x) have positive effect on the dependent variable (y)? Why or why not? Explain your answer.
登入後即可作答並保存紀錄。
此題考驗對多元線性迴歸分析結果的解讀能力,包括模型擬合、決定係數以及係數的顯著性與方向性。
核心觀念:
多元線性迴歸模型 嘗試找出自變數 () 與應變數 () 之間的線性關係。迴歸分析的輸出提供了模型擬合優劣的指標(如 )、各係數的顯著性(P-value)以及估計的係數值。
題目資訊:
- 應變數 : jet-turbine engine thrust
- 自變數:
- : primary speed of rotation
- : secondary speed of rotation
- : fuel flow rate
- 先驗知識:所有自變數對應變數皆有「正向」影響。
- 提供之報表為迴歸分析的輸出結果。
(a) What is the fitted linear regression model?
解答:
擬合的線性迴歸模型是由報表中的「Coefficients」部分給出的。模型的形式是 。
從表格中讀取估計的係數值:
- (Intercept) = 5763.816
- (x1) = 2.162198
- (x2) = 0.068896
- (x3) = -0.24037
因此,擬合的模型是:
(b) What is the coefficient of determination of this model?
解答:
係數決定係數 (Coefficient of Determination) 通常表示為 。它衡量了模型中自變數能解釋應變數變異的比例。
從「Regression Statistics」部分可以讀取:
解釋:
這個值表示模型中的自變數()解釋了應變數 總變異的約 99.32%。這是一個非常高的比例,表示模型對數據的擬合度非常好。
(c) Does the model coefficients consistent with the prior knowledge that all three independent variables (x) have positive effect on the dependent variable (y)? Why or why not? Explain your answer.
解答: