109 年 國立成功大學工業與資訊管理學系研究所乙組《統計學》
第 I-1 題6 分
It is known that two defective copies of a commercial software program were erroneously sent to a shipping lot that has now a total of 75 copies of the program. A sample of copies will be selected from the lot without replacement.
a) If three copies of the software are inspected, determine the probability that exactly one of the defective copies will be found. (3%)
b) If 73 copies are inspected, determine the probability that both copies will be found. (3%)
登入後即可作答並保存紀錄。
核心觀念
此題屬於不放回抽樣,且母體中只有兩份瑕疵軟體,因此使用超幾何分配。
若母體總數為 ,其中有 個瑕疵品,抽取 個樣本,令 表示抽到的瑕疵品數,則
本題資料為:
- 母體總數:
- 瑕疵品數:
- 正常品數:
解題方法
a) 抽查 3 份,恰好找到 1 份瑕疵品
令 為抽查 3 份中找到的瑕疵品數,要求 。
從 2 份瑕疵品中選出 1 份,再從 73 份正常品中選出 2 份,所有可能的 3 份樣本數為 ,因此
計算得
所以
亦即約為 。
b) 抽查 73 份,找到兩份瑕疵品
直接使用超幾何分配,要求在 73 份樣本中找到 2 份瑕疵品:
第 I-2 題9 分
The probability that your call to a service line is answered in less than 30 seconds is 0.75. Assume that your calls are independent.
a) If you call 20 times, what is the probability that at least 18 calls are answered in less than 30 seconds? (3%)
b) What is the probability that you must call four times to obtain the first answer in less than 30 seconds? (3%)
c) What is the probability that you must call six times in order for two of your calls to be answered in less than 30 seconds? (3%)
登入後即可作答並保存紀錄。
核心觀念
- 二項分布:在獨立、相同成功機率 的 次伯努利試驗中,成功次數 ,其機率質量函式
- 幾何分布:第一次成功出現在第 次試驗時,,機率質量函式
- 負二項分布(或稱「第 次成功的試驗次數」):第 次成功出現在第 次試驗時,,機率質量函式
本題中,成功=「通話在 30 秒內被接聽」,成功機率 ,且各次通話相互獨立。
a) 20 次通話,至少 18 次在 30 秒內被接聽
解題方法
以二項分布計算 ,需求
計算步驟
計算各項:
將數值代入(使用計算機或手算):
合計:
答案
【答案】(四捨五入至小數第三位)。
b) 必須撥打四次才能第一次在 30 秒內得到接聽
第 I-3 題10 分
A website gets approximately seven visitors per minute. Suppose the number of website visitors per minute follows a Poisson probability distribution.
a) What is the density function for the time between website visits? (5%)
b) What is the probability no one will access the website in a 12-second period? (5%)
登入後即可作答並保存紀錄。
核心觀念
網站每分鐘訪客數服從 Poisson 分配,表示訪客到達可視為速率固定的 Poisson process:
Poisson process 的相鄰兩次到達時間間隔 服從指數分配:
其機率密度函數為:
此外,長度為 的時間內沒有訪客到達,其機率為:
解題方法
(a)兩位訪客之間時間的密度函數
令 表示相鄰兩位訪客到達之間的時間,單位採用「分鐘」。
因為到達速率為每分鐘 人,所以:
因此, 服從參數為 的指數分配,其密度函數為:
當 時間間隔不可能為負,因此:
推導
在時間長度 內沒有訪客到達的機率為:
因此, 的累積分配函數為:
對 微分即可得到密度函數:
(b)12 秒內沒有訪客進入網站的機率
先統一時間單位:
第 I-4 題10 分
The mean travel time to work for individuals in a city is 31.5 minutes. Assume the population mean is μ = 31.5 minutes and the population standard deviation is σ = 12 minutes. A sample of 50 residents is selected.
a) What is the sampling distribution of the sample mean travel time to work for the 50 residents? (5%)
b) What is the probability that the sample mean will be within 3 minutes of the population mean? (5%)
登入後即可作答並保存紀錄。
本題考驗中央極限定理 (Central Limit Theorem, CLT) 的應用,以及樣本平均數的抽樣分佈。
核心觀念:
- 中央極限定理 (CLT):當從任何母體中抽取足夠大的樣本 (通常 ) 時,樣本平均數的抽樣分佈將近似於常態分佈,其平均值等於母體平均值,標準差(稱為標準誤)等於母體標準差除以樣本大小的平方根。
- 抽樣分佈 (Sampling Distribution):描述樣本統計量(在此為樣本平均數 )的機率分佈。
題目設定:
- 母體平均數 分鐘
- 母體標準差 分鐘
- 樣本大小
a) 求 50 位居民的樣本平均通勤時間的抽樣分佈。
由於樣本大小 ,大於 30,我們可以應用中央極限定理。
根據 CLT,樣本平均數 的抽樣分佈將近似於常態分佈。
-
平均值 (Mean):樣本平均數分佈的平均值等於母體平均值。
分鐘。 -
標準差 (Standard Deviation):樣本平均數分佈的標準差,稱為標準誤 (Standard Error, SE),計算公式為 。
分鐘。
因此,樣本平均數 的抽樣分佈近似於一個平均值為 31.5 分鐘,標準差約為 1.697 分鐘的常態分佈。
我們可以寫成 或 。
【答案】樣本平均數 的抽樣分佈近似於常態分佈,其平均值為 31.5 分鐘,標準差(標準誤)為 分鐘。
第 I-5 題15 分
A research group is interested in testing a manufacturer's claim that its new model car will travel at least 25 miles per gallon of gasoline. Assume that σ is 3 miles per gallon.
a) With α = 0.02 and a sample of 30 cars, what is the rejection rule based on the value of x̄ for the test to determine whether the manufacturer's claim should be rejected? (5%)
b) What is the probability of making a Type II error if the actual mileage is 23 mile? (5%)
c) What sample size would be recommended if the researcher wants an 80% chance of detecting that mean mileage is less than 25 miles per gallon when it actually 24? (5%)
登入後即可作答並保存紀錄。
核心觀念
- 單樣本平均數之母體標準差已知,使用 z 檢定。
- 原假設 (製造商聲稱),備擇假設 (欲檢驗是否低於 25 mpg)。
- 為第一類錯誤機率(左尾), 為第二類錯誤機率。
- 樣本均值的抽樣分佈:
其中 , 為樣本數。 - 臨界值由標準常態分位點 決定。
a) 拒絕規則
以 的左尾檢定:
-
取標準常態分位點
-
由
( 為原假設臨界均值)
拒絕域為 ,即
-
解得
拒絕規則:若樣本平均 mpg,則拒絕原假設,認為製造商的「至少 25 mpg」聲稱不成立。
b) 第二類錯誤機率()
當真實平均 時,未拒絕 的機率即 :
-
臨界均值 (前述結果)。
-
在 下,
第 II-1 題25 分
A two-factor experiment is designed to determine whether the operation times between two machines and those among three approaches are different and whether the interaction between these two factors exists. The data of the operation times are summarized in the following table.
| Approach 1 | Approach 2 | Approach 3 | |
|---|---|---|---|
| Machine 1 | 8 | 15 | 16 |
| Machine 2 | 6 | 12 | 9 |
Note: Each cell has 4 observations. (This information is crucial and was inferred from the calculation of degrees of freedom in a typical ANOVA table. The original problem might have omitted this detail or it's implied by the context of ANOVA).
Let's assume each cell has observations for the purpose of completing the ANOVA table.
a) (6%) State the hypotheses tested in this experiment.
b) (15%) Complete an ANOVA table.
c) (4%) Use α = 0.05, explicitly present your conclusions.
登入後即可作答並保存紀錄。
核心觀念
本題屬於兩因子完全隨機設計(Two‑factor completely randomised design)的單向變異數分析(One‑way ANOVA)。
- 因子 A(Machine):2 個水準
- 因子 B(Approach):3 個水準
- 每個處理組(cell) 有 次觀測。
分析目標:
- 檢驗 Machine 主效應是否顯著。
- 檢驗 Approach 主效應是否顯著。
- 檢驗 Machine × Approach 交互作用是否顯著。
解題方法
- 先把題目給的每格數值視為該格的 平均值,因每格有 個觀測,可把每格的總和寫成 。
- 計算 總平方和(SST)、因素 A、B 及交互作用的平方和(SSA、SSB、SSAB),再求 誤差平方和(SSE)。
- 由平方和除以相對自由度得到 均方(MS),再以 判斷顯著性(若 , 會趨於無窮大,即必定拒絕 )。
下面依步驟列出所需計算。
a) 假設陳述
| 因子 | 零假設 | 對立假設 |
|---|---|---|
| Machine(A) | (兩台機器的平均作業時間相等) | 至少有一台機器的平均作業時間不同 |
| Approach(B) | (三種作業方式的平均時間相等) | 至少有一種作業方式的平均時間不同 |
| 交互作用(A×B) | (無交互作用,模型為加法) | 存在交互作用(兩因子之效應非純加法) |
b) 完成 ANOVA 表
先列出全部必要的總和():
| Approach 1 | Approach 2 | Approach 3 | 小計 | |
|---|---|---|---|---|
| Machine 1 | ||||
| Machine 2 | ||||
| 小計 | (Grand Total ) |
- 總樣本數
- 總平方和
- 主效應 A(Machine)
- 主效應 B(Approach)
第 II-2 題8 分
A regression analysis of the relationship between the value of a factor (in thousands) and the sales price (in thousands of dollars) is conducted, as follows:
Analysis of Variance
SOURCE DF SS
Regression 1 310.2774
Error 5
Total 6 1002
Predictor Coef SE Coef
Constant 29.39911 4.807253
X 1.547478 0.463499
a) (3%) What is the value of the coefficient of determination?
b) (6%) Use the F statistic to test the significance of the relationship at a 0.05 level of significance.
c) (2%) Predict the sales price when the factor is at the value of 60,000.
登入後即可作答並保存紀錄。
本題考驗簡單線性迴歸 (Simple Linear Regression) 的相關概念,特別是變異數分析表 (ANOVA table) 的解讀、判定係數 (Coefficient of Determination) 的計算、迴歸模型的顯著性檢定,以及預測。
核心觀念:
- 判定係數 ():衡量迴歸模型解釋應變數變異的比例。
其中 SSR 是迴歸平方和,SSE 是誤差平方和,SST 是總平方和。 - 變異數分析表 (ANOVA Table):用於檢定迴歸模型的整體顯著性。
- (Regression Sum of Squares)
- (Error Sum of Squares)
- (Total Sum of Squares)
- -statistic
- 迴歸係數的顯著性檢定:通常使用 F 檢定來檢定迴歸模型的整體顯著性。對於簡單線性迴歸,F 檢定等同於檢定斜率係數 是否顯著不為零。
- 預測:使用迴歸方程式來預測應變數的值。
數據整理與初步計算:
從 ANOVA 表和 Predictor 表中,我們得到以下資訊:
- (Constant Coef)
- (X Coef)
a) 判定係數的值為何?
判定係數 可以通過以下公式計算:
我們也可以先計算 SSE,然後用 。
兩種方法計算結果非常接近,由於 SSR 和 SST 是直接給定的,使用 應更直接。
【答案】約 0.310
b) 使用 F 統計量檢定關係的顯著性,顯著水準為 0.05。
首先,我們需要完成 ANOVA 表,計算 F 統計量。
計算平均平方和 (MS) 和 F 統計量:
- -statistic
第 II-3 題12 分
A sample of 600 sales showed 270 for Product A, 230 for Product B, and 100 for other products. We would like to know whether the market shares of Products A and B are equal.
a) (4%) State the reason of the method you adopt for this problem?
b) (4%) What are the null and alternative hypotheses?
c) (6%) At a 0.05 level of significance, explicitly present your test procedure and conclusion?
登入後即可作答並保存紀錄。
核心觀念
- 本題屬於 分類資料的比例檢定,觀測值為多項分配 (multinomial) 的類別次數。
- 檢驗目標是比較 兩個類別 (A 與 B) 的母體比例是否相等,常用 卡方適合度檢定(或等價的兩比例 Z 檢定)於大樣本情形。
- 其主要假設、統計量與自由度如下:
- 觀察次數 ,期望次數 由 合併樣本比例 估計。
- 卡方統計量 ,自由度 。
解題方法
-
採用卡方適合度檢定的原因
- 觀測資料為 離散類別計數,且樣本大小 足以滿足近似正態條件(每類期望次數 )。
- 只關心兩個類別的比例是否相等,其他類別可視為「餘下」對檢驗結果不構成影響。
- 卡方檢定直接以 觀測次數 vs. 期望次數 之差異衡量,計算簡便且與兩比例 Z 檢定等價。
-
建構檢驗統計量
- 總樣本數 。
- 觀測次數:。
- 在 下,,以合併樣本估計 :
- 期望次數(以 為基礎,僅針對 A、B 兩類)
-
計算卡方統計量