108 年 國立清華大學動力機械工程學系碩士班丙組《統計學》
第 1 題30 分
- (30 pts.) True or False. 答對得3分,空白得0分,答錯得-2分(也就是,倒扣 2 分,所以請勿猜答案)(扣到本大題0分為止)
(a) T F The event-probability , where is any random variable and is any real value.
(b) T F We call an unbiased estimator of if .
(c) T F Suppose that follows a distribution with degrees of freedom 10 and follows a standard normal distribution. .
(d) T F If are random variables with the same mean and same variance , and with the lag-one correlation . The variance of the sample mean .
(e) T F Suppose that follows a distribution with degrees of freedom 17. , where . (Hint. See distribution table attached).
(f) T F Suppose that follows a Normal distribution with mean 0 and variance 1. , where . (Hint. See normal distribution table attached).
(g) T F Suppose that we obtain a 95% confidence interval of the mean to be . We know that .
(h) T F Let be a random sample of size . Let the sample mean be denoted as . The Central Limit Theorem is: , as .
(i) T F The terms ANOVA stands for "analysis of variance". In experimental design with two factors A and B, ANOVA is used to test the equality of variances of random variables, where each random variable denotes for responses for some factor (say Factor A) with a corresponding level.
(j) T F Let denote for (the population mean of one distribution) or (the difference of two population means of two associated distributions). Let be the estimator of . Suppose that the standard error of , denoted as , is known. The 95% confidence interval of the population parameter can be written as , where is some constant.
登入後即可作答並保存紀錄。
核心觀念
本題綜合考查:
- 隨機變數與事件機率
- 不偏估計量與變異數
- 分配與標準常態分配的尾端機率
- 樣本平均數的變異數與相關性
- 信賴區間的正確解釋
- 中央極限定理
- ANOVA 的用途
- 信賴區間的一般形式
(a)
敘述: 對任意隨機變數 及任意實數 ,皆有 。
判斷:錯(F)
連續型隨機變數通常滿足
但離散型隨機變數不一定。例如若 為擲骰子的點數,則
題目說的是「任意隨機變數」,只要存在離散型反例,敘述即錯誤。
解題技巧: 看到「任意隨機變數」時,應立即檢查離散型與連續型兩種情形,不能只套用連續型分配的性質。
(b)
敘述: 若 ,則稱 為 的不偏估計量。
判斷:錯(F)
不偏估計量的定義是
而
只表示 幾乎必定為某個固定常數,並不保證該常數等於 。
例如令 ,則
但若 ,則
因此 並非不偏估計量。
解題技巧:
「不偏」看期望值;「穩定程度」看變異數。兩者是不同概念。
(c)
敘述: 若 ,,則
判斷:對(T)
自由度為 的 分配相較於標準常態分配具有較厚的尾端,因此在正的臨界值 右側,
由表可得約為
而
因此確實有
解題技巧: 分配自由度有限時尾端較厚;當自由度增加, 分配逐漸接近標準常態分配。
(d)
敘述: 具有相同平均數與變異數,且相鄰變數的相關係數為 ,則
判斷:錯(F)
樣本平均數為
其變異數為
只有在所有 彼此不相關,亦即所有共變異數皆為 時,才有
本題給出相鄰變數的相關係數為 ,因此
這些共變異數會增加樣本平均數的變異數,所以不能直接寫成 。
解題技巧: 公式 的前提是獨立同分布,不能只知道平均數與變異數相同。
(e)
敘述: 若 ,且
則 。
判斷:對(T)
條件
表示 是自由度 的 分配之第 百分位數,即
查 分配表可得
因此
解題技巧: 右尾機率為 ,要查的是 ,不是左尾 的負值。
(f)
敘述: 若 ,則
且題目給出的區間為 。
判斷:錯(F)
題目中的區間變數應為 ,依條件
可知 是標準常態分配的第 百分位數。因此
所以正確範圍應為
第 2 題10 分
- (10 pts.) Consider a stochastic network, shown in Figure 1, with only one work-station (WS1), which has random capacity, say , determined by a known discrete probability distribution. Moreover, with positive probability , an output entity from WS1 is non-defective, . Let be the available items (or work in process, WIP) at the beginning of the process. Moreover, denotes the no. of items that allowed to enter WS1, and denotes the no. of non-defective items produced from WS1. Specifically,
• .
• for the above-mentioned one workstation network.
🖼️【此處有附圖,請對照原卷】Figure 1: A network with one workstation, WS1
The goal is to compute the probability that the random non-defective output of the above-mentioned network WS1 meets the pre-determined constant demand , i.e., .
Suppose that parameters used are , , and . The possible capacity values are given in Table 1.
Table 1: : Capacity of WS1
| X₁ | 0 | 1 | 2 | 3 | 4 | 5 | Others |
|---|---|---|---|---|---|---|---|
| prob. | 0.001 | 0.019 | 0.03 | 0.05 | 0.1 | 0.8 | 0 |
Questions.
(a) (5 pts.) Compute the probability , where .
(b) (5 pts.) Compute for the above-mentioned WS1.
登入後即可作答並保存紀錄。
本題考驗機率計算與隨機變數的定義。
(a) 首先計算 。
。當 時,。
我們需要計算 。
表示 。這意味著 必須是 4。
從 Table 1 中,我們知道 。
因此,。
接著計算 。
是非缺陷產品的數量。已知 是生產非缺陷產品的機率。
的產生是基於進入 WS1 的項目數量 進行二項分配的。
具體來說,。
我們需要計算 。
這表示在有 4 個進入的項目,且每個項目有 0.9 的機率是非缺陷的情況下,恰好有 4 個是非缺陷產品的機率。
根據二項分配的機率質量函數:
其中 , , 。
。
最後,計算兩個機率的乘積:
。
(b) 計算 。
且 。所以我們需要計算 。
是非缺陷產品的數量,其分佈取決於 。
。
。
我們需要考慮所有可能的 值,並根據其對應的機率進行加權平均。
的可能值取決於 的值:
- 若 ,
- 若 ,
- 若 ,
- 若 ,
- 若 ,
- 若 ,
- 若 , 。由於 是離散分佈,我們假設 "Others" 代表 的情況。如果 ,則 。
我們需要計算 。這意味著 可以是 4 或 5 (因為 的最大值是 5)。
我們將使用全機率公式,對 的所有可能值進行分類計算:
第 3 題10 分
- (10 pts.) Let be the size () (unit: cm) of some electrical component. To estimate the expected value of the size (), we take a random sample of size . Let be the sample mean. Assume that data are independent with the standard deviation () 0.3 cm.
(a) (5 pts.) Compute for . List any assumption if needed.
(b) (5 pts.) Find the smallest sample size such that . List any assumption if needed.
登入後即可作答並保存紀錄。
本題考驗中央極限定理的應用,以及信賴區間與樣本數的計算。
題目已知:
母體標準差 cm。
樣本大小 。
樣本平均數 。
母體平均數 。
的抽樣分配近似於常態分配,其期望值 。
的標準差(標準誤)為 。
(a) 計算 for 。
首先,我們計算樣本平均數的標準誤:
。
我們要計算的機率是 。
這可以改寫成 。
將不等式兩邊同時除以標準誤 ,我們得到標準常態變數 :
。
計算臨界值:
。
所以,我們要計算 。
這個機率等於 。
由於標準常態分配是對稱的,。
所以,機率等於 。
查標準常態分配表,當 時,。
當 時,。
我們可以近似取 (透過內插法或直接使用統計軟體)。
第 4 題30 分
- (30 points)
When diagnostics indicate that some assumptions of the linear regression model are violated, remedial measures may need to be taken. The assumption of common variance of the error term plays a key role in least squares. An approach to handling heterogeneous variances, called heteroscedasticity, is the use of weighted least squares.
Consider the simple linear regression model
, ,
where the errors are pairwise independent with mean 0, and the variance of is for some non-null coefficients that are not all equal (we have for some ). Since coefficients are not all equal, there is heteroscedasticity.
Consider , . The weighted least squares method is a generalization of the least squares that accounts for heteroscedasticity.
(10 points) (a) The weighted least squares estimators and of and , respectively, minimize the weighted least squares criterion
.
Prove or disprove the weighted least squares estimators
Show all your work. Partial work will not receive full credit.
(10 points) (b) Are the weighted least squares estimators and unbiased estimators of and , respectively? Justify and show all your work.
(10 points) (c) Show that the weighted least squares estimators are equal to the ordinary (unweighted) least squares estimators, i.e.,
when the errors have common variance . Show all the details of your proof.
登入後即可作答並保存紀錄。
核心觀念
本題考查異質變異下的加權最小平方法(weighted least squares, WLS)。
給定
且
令
加權最小平方法藉由最小化
使變異數較大的觀測值獲得較小權重,變異數較小的觀測值獲得較大權重。
定義加權平均數
(a) 推導加權最小平方法估計量
第一步:建立一階條件
令
則
分別對 與 微分:
令兩個偏導數皆為零,得到加權常態方程:
第二步:先解出
將估計量代入第一條常態方程:
因此
即
第三步:解出
將
代入第二條常態方程:
由於
整理後可得
因此題目所給的估計量正確:
以及
也可以使用加權離均差形式表示:
(b) 是否為不偏估計量
假設 視為固定值。若 為隨機變數,以下結果可解讀為條件於所有 下的不偏性。
的不偏性
由模型
代入加權離均差形式:
其中
因此
分子可分成兩部分:
注意
所以
因此
取期望:
第 5 題20 分
- (20 points)
(10 points) (a) Derive the expected mean squares of error in the two-way ANOVA table below, from the model
, , , .
| Term | SS |
|---|---|
| A | |
| B | |
| AB | |
| Error |
(10 points) (b) Prove or disprove that sum of squares of total equals sum of squares of regression plus sum of squares of error . (Here, . )
• : the overall average of the , i.e., .
• : the average of all responses for the -th A factor and -th B factor, i.e., .
• : the average of all responses for the -th A factor, i.e., .
• : the average of all responses for the -th B factor, i.e., .
• For each factor combination , the random error terms are ; the variance is the same for each factor combination.
• The random error terms are independent.
• Sum of squares of total .
• Mean squares of error .
登入後即可作答並保存紀錄。
本題考驗兩因子變異數分析 (Two-way ANOVA) 的期望均方 (Expected Mean Squares, EMS) 推導,以及總平方和 (SST) 的分解性質。
模型為:
其中 且獨立。
總共有 個 A 層級, 個 B 層級, 個重複觀測。總觀測數 。
(a) 推導 。
。
首先,我們需要計算 SSE 的期望值 ,並確定誤差的自由度。
將模型代入 SSE:
其中 。
所以,。
現在計算 :
由於 ,所以 。
由於 之間是獨立的,當 時,。
所以,。
。
。
由於 獨立, 。
所以,。
因此,。
將這些代回 的計算:
總共有 項。
。
誤差的自由度 (degree of freedom of error) 是總觀測數減去模型中所有參數的估計數。
在固效模型 (fixed effects model) 中,參數包括 , 個 , 個 , 個 。
總參數數為 。
自由度為 。
然而,這裡的 SSE 公式是針對 ,這暗示我們是基於因子組合的平均值來定義 SSE。
在 ANOVA 中,SSE 的自由度計算是:總觀測數 減去模型中被估計的參數數量。
在固效模型下,我們估計 。
總共有 個獨立的參數(如果考慮約束條件 for each i, for each j)。
或者,如果我們不考慮約束條件,而是估計 個 , 個 , 個 ,加上 ,總共有 個參數。
但通常我們是以約束條件來定義參數的獨立性。
對於 ,
若 , , for each i, for each j。
則獨立參數個數為 。
總觀測數為 。
自由度為 。
然而,SSE 的定義是基於 ,這意味著我們已經處理了 A 和 B 的效應。
。
這裡 是一個估計,它包含了 的資訊。
對於每一個 組合,有 個觀測值 。