109 年 國立成功大學交通管理研究所甲組《統計學》
第 1 題
- Long questions.
For 114 restaurants in NYC we have Zagat ratings on food as well as information on the price of a meal. We want to understand if there is a relationship between quality (as summarized by the Zagat ratings, the overall ratings are on a 30 point scale.) and price. We a run a simple regression of price on quality and obtain the following output:
| Estimate | Std error | t-ratio | P-value |
|---|---|---|---|
| Intercept | -18.154 | 6.553 | 2.77 |
| slope | 2.6253 | 0.3315 | 7.92 |
In addition, we know that and .
(a) Give an economic interpretation for the sign of the slope.
(b) Find an approximate 95% confidence interval for the true slope.
(c) Test the hypothesis that the slope is equal to zero at the 5% level. (You should be very precise here.)
(d) Test the hypothesis that the slope is equal to 2 at the 5% level. (You should use our usual rule-of-thumb here.)
(e) You are planning to have dinner in Soho at the restaurant YumYum. You know that the Zagat rating is 24. How much do you expect to pay?
(f) Find a 95% prediction interval for the total price given a rating equal to 24.
(g) Please interpret the meaning of the prediction interval in (f).
(h) Do you think quality of food is sufficient to explain the price of a meal? Briefly explain your answer.
登入後即可作答並保存紀錄。
核心觀念
本題考查一元線性迴歸:
其中:
- :截距,品質評分為 時的預期價格。
- :斜率,品質評分每增加 分,平均價格的預期變化量。
- 斜率估計值為 ,標準誤為 。
- 樣本數為 ,因此迴歸檢定自由度為
斜率的信賴區間與假設檢定皆使用 分配。
解題方法與各小題詳解
(a) 斜率正負號的經濟意義
估計斜率為
因此,餐廳的 Zagat 食物評分越高,預期餐費越高。
具體而言,在其他未納入模型的因素平均不變下,食物評分每增加 分,平均餐費約增加 美元。
正斜率表示品質與價格之間具有正向線性關係,但不代表品質提高必然直接造成價格上升,因為迴歸結果本身主要表示相關關係。
(b) 求真實斜率的近似 95% 信賴區間
95% 信賴區間為
因為 ,所以
計算誤差幅度:
因此:
所以近似 95% 信賴區間為
解釋為:在 95% 信賴水準下,真實斜率 約介於 與 之間。
(c) 在 5% 顯著水準下檢定斜率是否等於 0
假設為:
檢定統計量為
在雙尾檢定、、 下,臨界值約為
由於
故拒絕虛無假設。
題目表中的 值顯示為 ,但統計軟體顯示 並不代表真正的 ,而是表示 值小於顯示精度,實際上可寫為
因此,在 5% 顯著水準下,斜率顯著不等於 ,有充分統計證據支持品質評分與餐費之間存在線性關係。
(d) 在 5% 顯著水準下檢定斜率是否等於 2
假設為:
檢定統計量為
依題目要求使用通常的經驗法則:在 5% 雙尾檢定下,若
則拒絕虛無假設。
本題
所以不拒絕 。
結論是:在 5% 顯著水準下,沒有足夠證據認為真實斜率不同於 。因此,斜率等於 與資料並不矛盾。
(e) Zagat 評分為 24 時的預期餐費
迴歸方程式為
第 2 題
- A student analyzed data for a one-way analysis of variance situation for which there were 3 levels of the treatment, and 21 people measured at each level. Unfortunately, after running the analysis, the student lost the computer output. She said "All I remember is that one of the mean squares was 100 and the other one was 500, but I can't remember which was which. Oh, and I remember that the p-value for the test was about .01."
(a) Based on this information, can you construct the analysis of variance table? (The headings of the table structure are shown below to remind you.) If so, fill it in. If not, explain why not. If you think you can partially fill it in, do that.
| Source | SS | df | MS | F* | p-value |
|---|---|---|---|---|---|
| Treatment | |||||
| Error | |||||
| Total |
(b) In the statement of the question, it wasn't specified whether the treatment was fixed or random. Write the null and alternative hypotheses being tested in each of the two cases. Make sure you define any symbols you use.
(c) We suppose that the analysis of variance test is significant. Describe the procedure to determine which pairs of treatment are significantly different at 5% level. Justify your choice.
登入後即可作答並保存紀錄。
本題主要考驗考生對單因子變異數分析 (One-way ANOVA) 的理解,包括如何根據部分資訊重建 ANOVA 表、不同模型下的假設檢定,以及事後檢定 (post-hoc test) 的應用。
核心觀念:
- 單因子變異數分析 (One-way ANOVA) 的基本架構與公式
- 組間變異 (Between-group variation) 與組內變異 (Within-group variation)
- 均方 (Mean Square, MS) 的計算:
- F 統計量的計算:
- 假設檢定與 p-value 的關係
- 固定效應模型 (Fixed Effects Model) vs. 隨機效應模型 (Random Effects Model)
- 事後檢定 (Post-hoc tests)
解題過程:
題目給定資訊:
- 處理水準數 (Number of treatment levels) = 3
- 每組實驗單元數 (Number of experimental units per level) = 21
- 總實驗單元數
- 兩個均方 (MS) 的值為 100 和 500,但不知哪個是組間 (Treatment) 哪個是組內 (Error)。
- ANOVA 檢定的 p-value 約為 0.01。
自由度 (df) 計算:
- 組間自由度
- 總自由度
- 組內自由度
(a) Construct the analysis of variance table
我們知道 F 統計量是組間均方 (MS Treatment) 除以組內均方 (MS Error)。
。
我們也知道 p-value 約為 0.01。對於 ANOVA,p-value 是基於 F 分配計算的。
由於 p-value 很小 (0.01),這表示 F 統計量應該相對較大,因此分子 (MS Treatment) 應該大於分母 (MS Error)。
所以,我們可以推斷:
F 統計量 。
現在我們可以計算 SS (Sum of Squares):
現在我們可以填寫 ANOVA 表:
| Source | SS | df | MS | F* | p-value |
|---|---|---|---|---|---|
| Treatment | 1000 | 2 | 500 | 5.0 | 0.01 |
| Error | 6000 | 60 | 100 | ||
| Total | 7000 | 62 |
驗證:
我們計算得到的 F 統計量為 5。在 下,查 F 分配表,當 F=5 時,p-value 大約在 0.01 附近 (例如 約為 5.09)。這與題目給的 p-value 約 0.01 相符。
因此,這個 ANOVA 表是合理的。
【答案】
| Source | SS | df | MS | F* | p-value |
|---|---|---|---|---|---|
| Treatment | 1000 | 2 | 500 | 5.0 | 0.01 |
| Error | 6000 | 60 | 100 | ||
| Total | 7000 | 62 |
(b) Null and alternative hypotheses for fixed and random effects models
- 固定效應模型 (Fixed Effects Model):
在這個模型中,我們假設這三個處理水準是我們感興趣的全部水準,並且我們希望推論僅限於這三個水準。
設 為第 個處理水準的平均值。- 虛無假設 (): 所有處理水準的平均值都相等。
- 對立假設 (): 至少有一個處理水準的平均值與其他不同。
- 虛無假設 (): 所有處理水準的平均值都相等。
第 3 題
- A course can be taken for credit either by attending lecture sessions at fixed times and days, or by doing online session that can be done at the student's own pace and at those times the student chooses. The course coordinator wants to determine if these two ways of taking the course resulted in a significant difference in achievement as measured by the final exam for the course. The table below gives the scores on an examination with 45 possible points for one group of students and a second group of students who took the course with conventional lectures. The sample standard deviations of the two groups are and , respectively. We want to determine, via a one-sided hypothesis test, whether these data present sufficient evidence to indicate that the average grade for students who take the course online is significantly higher than those who attend a conventional class. (Given , )
| Group 1: online | 32 | 37 | 35 | 28 | 41 | 44 | 35 | 31 |
|---|---|---|---|---|---|---|---|---|
| Group 2: classroom | 35 | 31 | 29 | 25 | 34 | 40 | 27 | 32 |
(a) Develop appropriate hypotheses for the test.
(b) Compute the appropriate test statistic.
(c) What is the rejection region for the null hypothesis, with a significance level ? With the above test statistic, do we reject or not reject ?
(d) What is the one-sided confidence interval with a significance level ? Do we reject or not reject ?
(e) Find the p-value of the test statistic. With a significance level , do we reject or not reject ?
登入後即可作答並保存紀錄。
核心觀念
本題比較兩個獨立母體的平均數:
- 母體 1:採線上方式修課,平均成績為
- 母體 2:採傳統教室方式修課,平均成績為
欲檢定線上修課學生的平均成績是否較高,因此使用「兩獨立樣本平均數差」的右尾檢定。
由於兩組樣本數皆為 ,且題目提供兩組樣本標準差,可採兩獨立樣本 檢定。以不假設兩母體變異數完全相同的 Welch 檢定最為穩健。
題目表格中 Group 1 僅列出 筆資料,但題目明確給定 。依樣本數與給定的 ,可視為排版漏列第 筆資料;該筆應為 。因此線上組資料為:
解題方法與基本計算
線上組平均數為
教室組平均數為
因此樣本平均數差為
也就是線上組樣本平均成績高於教室組約 分。
兩平均數差的標準誤為
代入數值:
Welch 自由度為
實務上可取 或 ;兩者的臨界值與結論相同。
(a) 假設檢定
題目欲檢定線上修課學生的平均成績是否顯著較高,因此設:
也可將虛無假設寫成
而對立假設寫成
這是右尾、單尾檢定。
(b) 檢定統計量
檢定統計量為
代入計算:
因此檢定統計量為
(c) 拒絕域與檢定結論
顯著水準為
第 4 題
- Let and be two unbiased estimators for an unknown parameter . It is given that and . Furthermore, and are independent. Next, consider the estimator , where . For which value of has the smallest mean squared error (MSE)? (Provide the process of calculation.)
登入後即可作答並保存紀錄。
本題考驗考生對估計量均方誤差 (MSE) 的理解,以及如何利用獨立性與變異數性質來求解。
核心觀念:
- 無偏估計量 (Unbiased Estimator):。
- 均方誤差 (Mean Squared Error, MSE):。
- 獨立性 (Independence):若 X, Y 獨立,則 。
- 線性組合的期望值與變異數。
解題過程:
已知 和 是參數 的無偏估計量。
所以, 且 。
已知變異數: 且 。
和 是獨立的。
考慮新的估計量 ,其中 。
我們需要計算 的均方誤差 (MSE),並找出使 MSE 最小的 值。
1. 計算 的偏差 (Bias):
由於期望值是線性的:
因為 和 是無偏估計量,所以 且 。
。
所以,。
也是一個無偏估計量。
2. 計算 的變異數 (Variance):
由於 和 是獨立的,且 和 是常數:
第 5 題
- Consider the following ordered dataset: 0, 1, 3, 3, 4, 5, 5, 5, 7, 12
Which boxplot below corresponds to the given dataset? (Justify your answer.)
a.
b.
c.
d.
e. f.
🖼️【此處有附圖,請對照原卷】
登入後即可作答並保存紀錄。
核心觀念
盒鬚圖用五數摘要呈現資料:最小值、第一四分位數 、中位數 、第三四分位數 、最大值;若採用離群值判定,鬚端通常延伸到仍落在界限內的最遠資料點,超出界限的值另以點標示。
本題依圖形採用四分位數的線性插值法。四分位距與離群值界限為:
解題方法
資料共有 個,已由小到大排列:
中位數是第 、 個數的平均:
用線性插值求四分位數: 的位置是 ,介於第 個數 與第 個數 之間,因此 。 的位置是 ,介於第 個數 與第 個數 之間,因此 。
第 6 題
- Let X have a Exp(2) distribution. Using Chebyshev's inequality it holds that:
P (|X - “A” | < 2) ____ “B”, please find the number of A and B and also determine the inequality symbols to use (<, <, >, >) in the underline blank of the inequality. (Do NOT only show the results, it is necessary to provide the process of calculation.)
登入後即可作答並保存紀錄。
核心觀念
-
指數分配(Exponential Distribution)的母體參數:
在統計學標準定義中,若隨機變數 (以率參數 定義,機率密度函數 ),其期望值與變異數分別為: -
柴比雪夫不等式(Chebyshev's Inequality):
對任意具有有限期望值 與有限變異數 的隨機變數 ,以及任意常數 (或任意 ),其偏離平均數在 個標準差以內的機率下界滿足:或等價寫為(令 ):
解題方法
步驟一:求出指數分配 的期望值與變異數
題目給定 ,即率參數 :
標準差為:
步驟二:對應柴比雪夫不等式形式與求解
題目形式為 。
對照柴比雪夫不等式的基本形式 ,可知中心點 為隨機變數 的期望值 :
此時容許誤差 。
第 1.a 題2 分
(a) Give an economic interpretation for the sign of the slope.
登入後即可作答並保存紀錄。
本題為簡單線性迴歸模型的第一小題,要求解釋迴歸直線斜率的經濟意義。
核心觀念:
- 簡單線性迴歸模型:。
- 斜率 的解釋:表示自變數 X 每改變一個單位,應變數 Y 平均改變的量。
第 1.b 題3 分
(b) Find an approximate 95% confidence interval for the true slope.
登入後即可作答並保存紀錄。
本題要求計算迴歸直線斜率的 95% 信賴區間。
核心觀念:
- 迴歸係數的信賴區間:。
- t 分配的臨界值查找。
解題過程:
已知迴歸輸出結果:
- 估計的斜率
- 斜率的標準誤
- 樣本數
- 信賴水準為 95%,所以 。
1. 計算自由度 (df):
。
2. 查找 t 分配的臨界值:
對於 95% 信賴區間,我們需要 。
由於直接查表 可能不方便,我們可以使用近似值。
第 1.c 題3 分
(c) Test the hypothesis that the slope is equal to zero at the 5% level. (You should be very precise here.)
登入後即可作答並保存紀錄。
本題要求對迴歸直線的斜率進行假設檢定,檢定其是否顯著不為零。
核心觀念:
- 迴歸係數的假設檢定 (t-test)。
- 檢定統計量、臨界值、p-value 的應用。
解題過程:
我們要檢定虛無假設 對立於對立假設 。
顯著水準 。
樣本數 ,自由度 。
方法一:使用檢定統計量和臨界值
- 計算檢定統計量:
題目已給出斜率的 t-ratio,這就是我們的檢定統計量。
。
(題目輸出結果中的 t-ratio 就是 7.92,P-value 為 0.000)。
第 1.d 題3 分
(d) Test the hypothesis that the slope is equal to 2 at the 5% level. (You should use our usual rule-of-thumb here.)
登入後即可作答並保存紀錄。
本題要求對迴歸直線的斜率進行假設檢定,檢定其是否等於特定值 2。
核心觀念:
- 迴歸係數的假設檢定,當檢定值非零時。
- 比較檢定統計量與臨界值,或比較 p-value 與顯著水準。
解題過程:
我們要檢定虛無假設 對立於對立假設 。
顯著水準 。
樣本數 ,自由度 。
方法一:使用檢定統計量和臨界值
- 計算檢定統計量:
檢定統計量公式為:。
在此,,,。
。
第 1.e 題2 分
(e) You are planning to have dinner in Soho at the restaurant YumYum. You know that the Zagat rating is 24. How much do you expect to pay?
登入後即可作答並保存紀錄。
本題要求根據迴歸模型進行點預測 (point prediction)。
核心觀念:
- 迴歸模型的點預測:將自變數的已知值代入迴歸估計式,得到應變數的預測值。
解題過程:
我們已經建立了一個簡單線性迴歸模型,其估計式為:
其中 是預測的餐點價格 (Price),X 是 Zagat 評分 (Quality)。
從迴歸輸出結果得知:
第 1.f 題4 分
(f) Find a 95% prediction interval for the total price given a rating equal to 24.
登入後即可作答並保存紀錄。
本題要求計算一個新的觀察值(Zagat 評分等於 24)的預測區間 (prediction interval)。
核心觀念:
- 預測區間 (Prediction Interval) 的概念與公式。
- 預測區間比信賴區間 (Confidence Interval for the mean response) 更寬,因為它同時考慮了迴歸線的估計誤差和個別觀測值的隨機變異。
解題過程:
預測區間的公式為:
其中 是點預測值, 是臨界的 t 值,而 是預測值本身的標準誤。
的計算公式為:
其中 是殘差標準差, 是新的觀察值, 是樣本平均值, 是 X 的總平方和。
題目給定:
- 點預測值 (來自 1.e)。
- 殘差標準差 。
- 信賴水準 95%,,自由度 ,。
- 新的觀察值 。
- 樣本數 。
第 1.g 題4 分
(g) Please interpret the meaning of the prediction interval in (f).
登入後即可作答並保存紀錄。
本題要求解釋預測區間 (prediction interval) 的含義。
核心觀念:
- 預測區間的機率解釋。
- 與信賴區間 (confidence interval) 的區別。
解題過程:
在 (1.f) 中,我們計算了一個 95% 的預測區間為 ,用於預測 Zagat 評分為 24 的餐廳的餐點價格。
解釋:
一個 95% 的預測區間表示,在給定 Zagat 評分等於 24 的情況下,我們有 95% 的信心,該餐廳的實際餐點價格將落在我們計算出的區間 之間。
第 1.h 題4 分
(h) Do you think quality of food is sufficient to explain the price of a meal? Briefly explain your answer.
登入後即可作答並保存紀錄。
本題要求判斷食物品質 (Zagat 評分) 是否足以解釋餐點價格的變異。
核心觀念:
- 決定係數 () 的解釋。
- 代表模型能解釋的應變數變異的比例。
解題過程:
題目中給定的決定係數 (R-squared) 為 。
解釋:
決定係數 表示,在這個簡單線性迴歸模型中,Zagat 評分 (食物品質) 僅能解釋餐點價格變異的 35.9%。
第 2.a 題
(a) Based on this information, can you construct the analysis of variance table? (The headings of the table structure are shown below to remind you.) If so, fill it in. If not, explain why not. If you think you can partially fill it in, do that.
| Source | SS | df | MS | F* | p-value |
|---|---|---|---|---|---|
| Treatment | |||||
| Error | |||||
| Total |
登入後即可作答並保存紀錄。
本題要求根據部分提供的資訊,建構單因子變異數分析 (One-way ANOVA) 表。
核心觀念:
- 單因子 ANOVA 的結構:Source, SS, df, MS, F, p-value。
- MS 的計算:。
- F 統計量的計算:。
- 自由度 (df) 的計算:, , 。
- p-value 與 F 統計量的關係。
解題過程:
已知資訊:
- 處理水準數 (k) = 3
- 每組實驗單元數 = 21
- 總實驗單元數 (N) =
- 兩個均方 (MS) 的值為 100 和 500,但不知道哪個是組間 (Treatment) 哪個是組內 (Error)。
- ANOVA 檢定的 p-value 約為 0.01。
1. 計算自由度:
- 組間自由度 。
- 總自由度 。
- 組內自由度 。
2. 推斷 MS 值:
F 統計量是 。
p-value = 0.01 是一個較小的數值,這表示 F 統計量應該相對較大,因此分子 () 必須大於分母 ()。
所以,我們可以推斷:
第 2.b 題6 分
(b) In the statement of the question, it wasn't specified whether the treatment was fixed or random. Write the null and alternative hypotheses being tested in each of the two cases. Make sure you define any symbols you use.
登入後即可作答並保存紀錄。
本題要求針對單因子變異數分析中的兩種常見模型(固定效應模型和隨機效應模型),分別寫出其虛無假設和對立假設,並定義所使用的符號。
核心觀念:
- 固定效應模型 (Fixed Effects Model) vs. 隨機效應模型 (Random Effects Model)。
- ANOVA 的假設檢定。
解題過程:
1. 固定效應模型 (Fixed Effects Model):
在此模型下,我們假設這 k 個處理水準是所有我們感興趣的、固定且有限的。我們想比較這 k 個特定水準的平均值。
-
符號定義:
- : 第 個處理水準的母體平均值 ()。
-
虛無假設 (): 所有處理水準的母體平均值都相等。
-
對立假設 (): 至少有一個處理水準的母體平均值與其他不同。
。
(或者可以寫成: Not all are equal.)
2. 隨機效應模型 (Random Effects Model):
在此模型下,我們假設這 k 個處理水準是從一個更大的母體中隨機抽樣出來的。我們感興趣的是這個母體的平均值是否存在變異。
- 符號定義:
- : 整體母體平均值。
第 2.c 題6 分
(c) We suppose that the analysis of variance test is significant. Describe the procedure to determine which pairs of treatment are significantly different at 5% level. Justify your choice.
登入後即可作答並保存紀錄。
本題要求在 ANOVA 檢定結果顯著的前提下,描述如何進行事後檢定 (post-hoc test) 以找出哪些處理水準對之間存在顯著差異,並說明選擇特定方法的理由。
核心觀念:
- 事後檢定 (Post-hoc tests) 的目的與時機。
- 常見的事後檢定方法及其適用性。
- 比較多重比較下的錯誤率控制 (Family-wise Error Rate, FWER)。
解題過程:
由於 ANOVA 檢定結果顯著 (p=0.01 < =0.05),我們知道至少有一對處理的平均值是不同的。為了找出具體是哪一對或哪些對平均值有顯著差異,我們需要進行事後檢定。
1. 事後檢定的程序描述:
a. 確認 ANOVA 檢定結果顯著: 這是進行事後檢定的前提。
b. 選擇事後檢定方法: 根據研究目的、樣本數、變異數是否相等,選擇合適的檢定方法。
c. 執行檢定: 對所有感興趣的配對(或特定配對)進行檢定。
d. 比較結果: 根據各檢定的臨界值或 p-value,判斷哪些配對的差異是顯著的。
2. 選擇事後檢定方法與理由:
題目中,有 k=3 個處理水準,每組樣本數相等 ()。在這種情況下,最常用且推薦的事後檢定方法是 Tukey's Honestly Significant Difference (HSD) test。
- Tukey's HSD test 的優點:
- 比較所有配對: 它可以同時比較所有可能的配對平均值。
- 控制 FWER: 它能控制在進行多次配對比較時,至少出現一個第一類錯誤 (Type I error, 錯誤地拒絕為真的虛無假設) 的機率,即整體錯誤率 (Family-wise Error Rate, FWER)。
- 適用於樣本數相等: 在本例中,由於 ,Tukey's HSD test 是非常合適的。
- 計算公式: 。
- 是 Tukey's studentized range distribution 的臨界值。
- 是來自 ANOVA 表的組內均方。
第 3.a 題4 分
(a) Develop appropriate hypotheses for the test.
登入後即可作答並保存紀錄。
本題要求為比較兩組獨立樣本的平均數,建立適當的虛無假設 () 和對立假設 ()。
核心觀念:
- 獨立樣本 t 檢定 (Independent samples t-test) 的假設設定。
- 單尾檢定 (One-tailed test) 的設定,根據題目敘述判斷檢定方向。
解題過程:
題目要求檢定「average grade for students who take the course online is significantly higher than those who attend a conventional class」。
設:
- : online 群體的平均最終考試分數。
- : conventional class 群體的平均最終考試分數。
第 3.b 題4 分
(b) Compute the appropriate test statistic.
登入後即可作答並保存紀錄。
本題要求計算進行兩獨立樣本平均數比較時,適當的檢定統計量。
核心觀念:
- 獨立樣本 t 檢定統計量的計算。
- 變異數同質性檢定 (F-test for equality of variances) 的結果對 t 檢定公式的影響。
解題過程:
首先,我們需要計算兩個樣本的平均數和標準差。
數據集:
- Group 1 (online): 32, 37, 35, 28, 41, 44, 35, 31 (n1=9)
- Group 2 (classroom): 35, 31, 29, 25, 34, 40, 27, 32, 31 (n2=9)
給定的樣本標準差:
1. 進行變異數同質性檢定:
我們需要檢定 。
F 檢定統計量 。
自由度為 。
在 的雙尾檢定下,臨界值 。
第 3.c 題4 分
(c) What is the rejection region for the null hypothesis, with a significance level ? With the above test statistic, do we reject or not reject ?
登入後即可作答並保存紀錄。
本題要求確定單尾檢定的拒絕區域,並根據計算出的檢定統計量做出決策。
核心觀念:
- 單尾檢定的拒絕區域。
- 將檢定統計量與臨界值比較以做出決策。
解題過程:
根據 (3.a),我們的假設是:
這是右尾單一檢定 (right-tailed test)。
顯著水準 。
自由度 (根據 (3.b) 的計算)。
第 3.d 題4 分
(d) What is the one-sided confidence interval with a significance level ? Do we reject or not reject ?
登入後即可作答並保存紀錄。
本題要求計算單側信賴區間,並根據信賴區間的結果做出決策。
核心觀念:
- 單側信賴區間與單尾檢定的關係。
- 利用信賴區間判斷是否拒絕虛無假設。
解題過程:
我們的檢定是 vs 。
這是一個右尾檢定,對應的信賴區間是 下側信賴區間 (lower-sided confidence interval),其形式為 。
顯著水準 ,自由度 。
臨界值 。
標準誤 (根據 (3.b) 的計算)。
樣本平均數差 。
1. 計算單側信賴區間的上限 (U):
第 3.e 題4 分
(e) Find the p-value of the test statistic. With a significance level , do we reject or not reject ?
登入後即可作答並保存紀錄。
本題要求計算檢定統計量的 p-value,並根據 p-value 與顯著水準的比較來做出決策。
核心觀念:
- p-value 的定義。
- 單尾檢定的 p-value 計算。
- p-value 與顯著水準的比較規則。
解題過程:
我們進行的檢定是 ,這是一個右尾檢定。
檢定統計量為 (使用更精確的值),自由度 。
1. 計算 p-value:
p-value 是指在虛無假設為真的情況下,觀察到目前的檢定統計量值或更極端的結果的機率。
對於右尾檢定,p-value 。
p-value 。
由於 t 分配是對稱的,且 是一個非常接近 0 的負值:
由於對稱性,。
第 4.a 題
- Let and be two unbiased estimators for an unknown parameter . It is given that and . Furthermore, and are independent. Next, consider the estimator , where . For which value of has the smallest mean squared error (MSE)? (Provide the process of calculation.)
登入後即可作答並保存紀錄。
本題要求找到一個線性組合估計量 使其均方誤差 (MSE) 最小的參數 值。
核心觀念:
- 均方誤差 (MSE) 的定義:。
- 無偏估計量 (Bias = 0)。
- 獨立隨機變數的變異數性質: (若 X, Y 獨立)。
- 二次函數的最小值求解。
解題過程:
已知:
- 是 的無偏估計量,所以 且 。
- 且 。
- 和 是獨立的。
- ,其中 。
步驟 1:計算 的偏差 (Bias)。
由於期望值是線性的:
因為 是無偏估計量:
。
因此,。
也是一個無偏估計量。
步驟 2:計算 的變異數 (Variance)。
由於 和 是獨立的,且 和 是常數:
第 5 題
- Consider the following ordered dataset: 0, 1, 3, 3, 4, 5, 5, 5, 7, 12
Which boxplot below corresponds to the given dataset? (Justify your answer.)
a.
b.
c.
d.
e. f.
🖼️【此處有附圖,請對照原卷】
登入後即可作答並保存紀錄。
核心觀念
盒鬚圖用五數摘要呈現資料:最小值、第一四分位數 、中位數 、第三四分位數 、最大值;若採用離群值判定,鬚端通常延伸到仍落在界限內的最遠資料點,超出界限的值另以點標示。
本題依圖形採用四分位數的線性插值法。四分位距與離群值界限為:
解題方法
資料共有 個,已由小到大排列:
中位數是第 、 個數的平均:
用線性插值求四分位數: 的位置是 ,介於第 個數 與第 個數 之間,因此 。 的位置是 ,介於第 個數 與第 個數 之間,因此 。
第 6 題
- Let X have a Exp(2) distribution. Using Chebyshev's inequality it holds that:
P (|X - “A” | < 2) ____ “B”, please find the number of A and B and also determine the inequality symbols to use (<, <, >, >) in the underline blank of the inequality. (Do NOT only show the results, it is necessary to provide the process of calculation.)
登入後即可作答並保存紀錄。
核心觀念
-
指數分配(Exponential Distribution)的母體參數:
在統計學標準定義中,若隨機變數 (以率參數 定義,機率密度函數 ),其期望值與變異數分別為: -
柴比雪夫不等式(Chebyshev's Inequality):
對任意具有有限期望值 與有限變異數 的隨機變數 ,以及任意常數 (或任意 ),其偏離平均數在 個標準差以內的機率下界滿足:或等價寫為(令 ):
解題方法
步驟一:求出指數分配 的期望值與變異數
題目給定 ,即率參數 :
標準差為:
步驟二:對應柴比雪夫不等式形式與求解
題目形式為 。
對照柴比雪夫不等式的基本形式 ,可知中心點 為隨機變數 的期望值 :
此時容許誤差 。