112 年 國立臺北大學都市計劃研究所乙組《統計學》
第 1 題
一、True and False (是非題) questions. [State the reason for your answer. 3% for each, 30%]
- A distribution table is required when using the p value of a test statistic to make a decision.
登入後即可作答並保存紀錄。
這題考驗對假設檢定中 p-value 的理解。
p-value 的定義是:在虛無假說 (null hypothesis) 為真的情況下,觀察到目前樣本結果或更極端結果的機率。p-value 的計算通常是基於檢定統計量 (test statistic) 的分佈,而不是直接依賴於某種特定的「分佈表」。雖然在某些情況下,我們可能會查閱標準常態分佈表或 t 分佈表來找到 p-value 所對應的臨界值,但 p-value 本身就是一個機率值,其計算與決策不必然需要一個「分佈表」作為前提。
第 2 題
- A random variable cannot be negative.
登入後即可作答並保存紀錄。
這題考驗對隨機變數定義的理解。
隨機變數 (random variable) 是將隨機實驗的結果映射到實數的函數。隨機變數的值域 (range) 取決於隨機實驗本身。有些隨機變數的值域確實不能是負數,例如:
- 計數型的隨機變數(如:一天內到達商店的顧客數),其值只能是非負整數(0, 1, 2, ...)。
- 表示比例或機率的隨機變數,其值域在 [0, 1] 之間。
第 3 題
- Assuming a random variable X has a normal distribution with mean and standard deviation , then .
登入後即可作答並保存紀錄。
這題考驗對常態分佈 (normal distribution) 性質的理解。
常態分佈是一個對稱的分佈,其對稱中心就是它的平均數 。
對於任何一個對稱分佈,其平均數(或中位數)會將分佈的總機率 1 分成兩半,也就是說,落在平均數左邊(小於平均數)的機率是 0.5,落在平均數右邊(大於平均數)的機率也是 0.5。
具體來說,對於一個平均數為 的常態分佈 :
第 4 題
- Consider two events A and B and suppose , then .
登入後即可作答並保存紀錄。
這題考驗對條件機率 (conditional probability) 與事件獨立性 (independence) 的理解。
題目敘述為:「假設有兩事件 A 和 B,且 ,則 。」
是在事件 B 已經發生的條件下,事件 A 發生的機率。
是事件 A 發生的機率。
表示「事件 B 的發生,增加了事件 A 發生的機率」。換句話說,事件 B 的發生對事件 A 具有「正向關聯」或「提升」的作用。
然而,這個敘述並非總是成立的。事件 A 和 B 的關係有三種可能:
- A 的發生會增加 B 發生的機率 (A and B are positively associated): 此時 且 。
- A 的發生會降低 B 發生的機率 (A and B are negatively associated): 此時 且 。
- A 的發生與 B 的發生無關 (A and B are independent): 此時 且 。
題目中的敘述 只描述了第一種情況,而沒有考慮到其他兩種可能性。因此,這個陳述不一定為真。
舉例反例:
假設一個袋子裡有 10 個球,其中 5 個是紅球 (R),5 個是藍球 (B)。再假設有 3 個球是大的 (L),7 個球是小的 (S)。
假設:
- 3 個大球都是紅球 (L 且 R)。
第 5 題
- Consider two events A and B, then we have .
登入後即可作答並保存紀錄。
這題考驗對機率加法法則 (addition rule of probability) 的理解。
機率的加法法則指出:
其中, 是事件 A 或事件 B 至少發生一個的機率。
是事件 A 發生的機率。
是事件 B 發生的機率。
是事件 A 和事件 B 同時發生的機率。
由於機率值 永遠是非負的 (),因此:
第 6 題
- Events A and B are mutually exclusive which imply they are independent.
登入後即可作答並保存紀錄。
這題考驗對「互斥事件」與「獨立事件」定義的理解及其關係。
-
互斥事件 (Mutually Exclusive Events): 如果兩個事件 A 和 B 不能同時發生,則稱它們為互斥事件。這意味著它們的交集為空集合,即 。
-
獨立事件 (Independent Events): 如果事件 A 的發生與否不影響事件 B 發生的機率(反之亦然),則稱它們為獨立事件。數學上,這表示 。
題目聲稱「互斥事件必然是獨立事件」。讓我們來檢驗這個說法。
假設 A 和 B 是互斥事件,則 。
如果它們同時也是獨立事件,則必須滿足 。
所以,必須有 。
這意味著,至少有一個事件的機率必須為 0。
- 如果 ,那麼 A 是一個不可能事件。
- 如果 ,那麼 B 是一個不可能事件。
然而,在大多數實際情況中,我們考慮的事件都不是不可能事件,即 且 。
如果 且 ,則 。
第 7 題
- The probability of an event A has to be less than 1.
登入後即可作答並保存紀錄。
這題考驗對機率基本性質的理解。
機率的定義域要求任何事件的機率 都必須滿足:
其中:
- 表示事件 E 是一個不可能事件 (impossible event)。
- 表示事件 E 是一個必然事件 (certain event)。
第 8 題
- The shape of the binomial distribution is symmetry when the probability of success equals 0.5.
登入後即可作答並保存紀錄。
這題考驗對二項分佈 (binomial distribution) 形狀特徵的理解。
二項分佈 描述了在 次獨立的伯努利試驗中,成功次數 的機率。其中 是單次試驗成功的機率。
二項分佈的形狀取決於參數 和 :
- 當 時:
二項分佈是對稱的。成功和失敗的機率相等,因此分佈的機率質量函數 (PMF) 在 的兩側是對稱的。具體來說,。這時,分佈呈鐘形(如果 較大)或对称的條狀圖。
第 9 題
- The significance level of a test statistic has to be given in advance when performing the hypothesis test problem.
登入後即可作答並保存紀錄。
這題考驗對假設檢定 (hypothesis testing) 中顯著水準 (significance level, ) 的理解。
在進行假設檢定時,我們需要預先設定一個顯著水準 。這個 代表了我們願意承擔的犯第一類錯誤 (Type I error) 的最大機率。第一類錯誤是指當虛無假說 (null hypothesis, ) 為真時,卻拒絕了 。
常見的顯著水準有 0.05 (5%)、0.01 (1%) 或 0.10 (10%)。
為什麼需要預先設定 ?
- 客觀性: 如果在看到檢定結果(例如 p-value)之後才決定 的值,那麼這個決策就可能受到結果的影響,變得主觀。
第 10 題
- When the correlation of coefficient of random variables X and Y equals zero, we have X and Y are independent.
登入後即可作答並保存紀錄。
這題考驗對相關係數 (correlation coefficient) 與獨立性 (independence) 之間關係的理解。
相關係數 (或樣本相關係數 ) 衡量的是兩個隨機變數之間的線性關係強度和方向。
- 表示完美的線性正相關。
- 表示完美的線性負相關。
- 表示兩個變數之間沒有線性關係。
獨立性 (independence) 是一個比線性關係更強的概念。如果兩個隨機變數 X 和 Y 是獨立的,則它們之間不存在任何形式的關係,包括線性關係、非線性關係等。
關鍵點:
- 獨立 相關係數為 0: 如果 X 和 Y 是獨立的,那麼它們之間沒有線性關係,所以相關係數 必定為 0。
- 相關係數為 0 獨立: 如果 X 和 Y 的相關係數為 0,這只表示它們之間沒有「線性」關係。但它們之間可能存在「非線性」的關係。
反例:
考慮一個隨機變數 X,其值為 -1, 0, 1,各機率為 1/3。
考慮另一個隨機變數 Y,其值為 。
則 Y 的可能值為 , , 。
Y 的機率分佈為:;
第 1 題
二、Multiple choice (單選題) [2% for each, 70%]
- What is the data type of the floor area (unit: in square foot)?
A. Count
B. Interval
C. Nominal
D. Ordinal
E. Ratio
登入後即可作答並保存紀錄。
這題考驗對資料類型 (data types) 的基本認識。
資料類型主要分為四種尺度:名目尺度 (Nominal)、次序尺度 (Ordinal)、區間尺度 (Interval) 和比例尺度 (Ratio)。
- 名目尺度 (Nominal): 資料只能用來分類,沒有順序。例如:性別(男/女)、血型(A/B/AB/O)。
- 次序尺度 (Ordinal): 資料具有分類和順序,但等級之間的差距無法確定或不均等。例如:學歷(國中/高中/大學)、滿意度(非常不滿意/不滿意/普通/滿意/非常滿意)。
- 區間尺度 (Interval): 資料具有分類、順序,且等級之間的差距是固定的、有意義的,但沒有絕對的零點。例如:攝氏溫度(0°C 不是沒有溫度)、華氏溫度。
- 比例尺度 (Ratio): 資料具有分類、順序、固定差距,並且有一個絕對的零點。絕對零點表示「沒有」該屬性。比例尺度的資料可以進行乘除運算。例如:身高、體重、收入、長度、面積、數量。
第 2 題
- What is the data type of the type of transportation to work such as walk, bus, metro or car?
A. Count
B. Interval
C. Nominal
D. Ordinal
E. Ratio
登入後即可作答並保存紀錄。
這題考驗對資料類型的認識,特別是分類資料。
題目問的是「上班的交通方式」,例如:步行、公車、捷運、汽車。
這些選項是不同的類別。
- 名目尺度 (Nominal): 資料用來分類,沒有順序。例如:性別、血型。
- 次序尺度 (Ordinal): 資料有順序,但等級間的差距不確定。例如:滿意度(非常滿意、滿意、普通…)。
在「步行」、「公車」、「捷運」、「汽車」這些交通方式之間,我們可以將它們視為不同的類別。它們之間沒有內在的數學順序。
第 3 題
- What is the data type for the preference of a certain type of coffee measured in a five level likert scale (like very much, like, fair, dislike, very dislike)?
A. Count
B. Interval
C. Nominal
D. Ordinal
E. Ratio
登入後即可作答並保存紀錄。
這題考驗對李克特量表 (Likert scale) 資料類型的判斷。
李克特量表是一種常用的測量態度、意見或偏好的方法。它通常提供一系列敘述,並讓受訪者從一個等級量表中選擇一個最能代表其意見的選項。這個量表有明確的順序。
題目中的五點李克特量表(非常喜歡、喜歡、普通、不喜歡、非常不喜歡)具有以下特徵:
- 分類: 將受訪者的偏好歸類到五個等級中。
- 順序: 這些等級有明確的順序,從「非常喜歡」到「非常不喜歡」。
- 等級差距: 雖然等級之間有順序,但等級之間的「差距」是否相等是有爭議的。例如,「非常喜歡」到「喜歡」的差距,是否與「喜歡」到「普通」的差距相等,是無法確定的。通常我們假設它們相等,但這是一個簡化。
第 4 題
- Which of the following is the most appropriate statistic to summarize the type of transportation to work such as walk, bus, metro or car?
A. Interquartile
B. Mean
C. Median
D. Range
E. None of the above
登入後即可作答並保存紀錄。
核心觀念
「通勤方式」分為步行、公車、捷運、汽車等類別,屬於名目尺度資料。各類別只有名稱上的差異,沒有大小順序,也不能進行加減或計算平均數。
描述名目資料時,適合使用眾數,也就是出現次數最多的類別。題目選項中沒有眾數,因此應選「以上皆非」。
解題方法
先判斷資料的尺度,再確認統計量是否適用:
- 通勤方式是類別資料,沒有數值大小或先後次序。
- 平均數、中位數、全距與四分位距都需要數值資料,部分還需要資料具有順序。
- 眾數可用於名目資料,但不在選項中。
因此,選項中的統計量都不適合;最適當的答案是 E。
選項分析
- A. 四分位距(Interquartile):四分位距是第三四分位數減第一四分位數,用來描述數值資料中央一半的分散程度。
第 5 題
- Which of the following is the least appropriate statistical statistic to summary the central tendency of the data for the floor area (unit: square foot)?
A. Mean
B. Median
C. Midpoint
D. Midrange
E. Mode
登入後即可作答並保存紀錄。
這題考驗對集中趨勢 (central tendency) 統計量在不同資料類型上的適用性,特別是對於比例尺度資料。
題目問的是「樓地板面積」(單位:平方英尺)的集中趨勢,這是一種 比例尺度 (Ratio) 的資料。
我們來分析各個選項:
- A. Mean (平均數): 對於比例尺度資料,平均數是最常用且最能代表集中趨勢的統計量。
- B. Median (中位數): 對於比例尺度資料,中位數也是一個有效的集中趨勢指標,尤其是在資料有極端值(離群值)時,中位數比平均數更穩健。
- C. Midpoint (中點): 這個詞比較模糊,但通常指的是區間的中點,或者在某些情況下可能與中位數或平均數混淆。如果它指的是某個區間的中點,那它不是一個描述整個資料集集中趨勢的統計量。如果它指的是一個資料集的「中間值」,那它可能與中位數類似,但「Midpoint」本身作為一個標準統計量術語比較少見,除非在特定上下文中(例如:區間的中點)。
- D. Midrange (中距): 中距的計算方式是 (最小值 + 最大值) / 2。它描述了資料範圍的中心點。對於比例尺度資料,中距可以計算,但它非常容易受到極端值(離群值)的影響,不如中位數或平均數穩健。
- E. Mode (眾數): 眾數是指資料集中出現頻率最高的數值。對於連續型比例尺度資料(如樓地板面積,雖然理論上是連續的,但實際測量時可能是離散的),眾數可能不存在,或者可能有多個眾數,或者只有一個數值出現一次,無法代表集中趨勢。即使資料是離散的,眾數也只能代表最常見的值,不一定代表「中心」。
第 6 題
- Which of the following is the lease appropriate statistic to summary the variability of the data for the floor area (unit: square foot)?
A. Correlation
B. Interquartile
C. Range
D. Standard deviation
E. Variance
登入後即可作答並保存紀錄。
這題考驗對變異性 (variability) 統計量在不同資料類型上的適用性。
題目問的是「樓地板面積」(單位:平方英尺)的變異性。這是一種 比例尺度 (Ratio) 的資料。
我們需要找出「least appropriate」(最不恰當)的變異性統計量。
- A. Correlation (相關係數): 相關係數是用來衡量兩個變數之間線性關係的強度,而不是衡量單一變數的變異性。因此,它絕對不適用於描述單一變數的變異性。
- B. Interquartile (四分位距): 四分位距 (IQR = ) 是衡量資料中間 50% 的離散程度,對離群值穩健,適用於次序尺度、區間尺度和比例尺度資料。
- C. Range (全距): 全距 (Max - Min) 是衡量資料的總體離散程度,適用於次序尺度、區間尺度和比例尺度資料,但對離群值非常敏感。
第 7 題
- Which of the following is the most appropriate statistic to summary the association between the data for the floor area (unit: square foot) and household income (unit: thousand dollars)?
A. Correlation
B. Interquartile
C. Mean
D. Relative frequency
E. Skewness
登入後即可作答並保存紀錄。
這題考驗對衡量兩個變數之間關係的統計量選擇。
題目問的是「樓地板面積」(比例尺度)與「家庭收入」(比例尺度)之間的關聯性。
我們需要尋找最適合用來描述這兩個數值型變數之間關係的統計量。
- A. Correlation (相關係數): 相關係數(如 Pearson correlation coefficient)是衡量兩個數值型變數之間線性關係強度和方向的標準統計量。它非常適合用來描述樓地板面積和家庭收入之間的線性關聯。
- B. Interquartile (四分位距): 四分位距是衡量單一變數的變異性,不適用於衡量兩個變數之間的關聯。
第 8 題
- Which of the following is the most appropriate statistic to examine the shape of the distribution for the floor area (unit: square foot)?
A. Correlation
B. Interquartile
C. Mean
D. Relative frequency
E. Skewness
登入後即可作答並保存紀錄。
這題考驗對描述分佈形狀 (shape of distribution) 的統計量選擇。
題目問的是「樓地板面積」(比例尺度)的「分佈形狀」。
- A. Correlation (相關係數): 衡量兩個變數之間的線性關係,與單一變數的分佈形狀無關。
- B. Interquartile (四分位距): 衡量資料的離散程度,間接反映分佈的寬度,但不是直接描述形狀(如對稱性、偏態)。
- C. Mean (平均數): 衡量集中趨勢,不能描述分佈形狀。
第 9 題
- A survey is distributed to understand the commuting methods to NTPU. 270 students were responded and the data were summarized as follows:
| Level of students | Method of Commutes | ||
|---|---|---|---|
| Metro | Bus | Walk | |
| freshman | 70 | 20 | 65 |
| Sophomore | 50 | 40 | 25 |
Determine the probability of the students who commute to school using metro or bus.
A. 1/9
B. 2/9
C. 1/3
D. 4/9
E. 2/3
登入後即可作答並保存紀錄。
這題考驗對機率計算的理解,特別是「或」事件的機率。
題目要求計算「學生搭乘捷運 (Metro) 或公車 (Bus) 上學的機率」。
首先,我們需要從表格中匯總相關數據。
總學生人數 N = 270。
我們可以計算出搭乘捷運的總人數:
freshman 搭乘捷運 = 70
Sophomore 搭乘捷運 = 50
總搭乘捷運人數 = 70 + 50 = 120
我們可以計算出搭乘公車的總人數:
freshman 搭乘公車 = 20
Sophomore 搭乘公車 = 40
總搭乘公車人數 = 20 + 40 = 60
接下來,我們需要計算搭乘「捷運或公車」的總人數。
由於「搭乘捷運」和「搭乘公車」是兩個互斥的事件(一個學生不能同時以捷運和公車作為主要通勤方式),因此可以使用機率的加法法則:
或者,直接計算總人數:
搭乘捷運或公車的總人數 = ( freshman 捷運 + freshman 公車 ) + ( Sophomore 捷運 + Sophomore 公車 )
= (70 + 20) + (50 + 40)
= 90 + 90
= 180
第 10 題
- (Continued Question 9.) Given the respondents who were freshman, determine the probability of the students who walked to school?
A. 1/6
B. 13/31
C. 1/3
D. 5/18
E. 13/18
登入後即可作答並保存紀錄。
這題是上一題的延續,考驗條件機率 (conditional probability) 的計算。
題目要求在「受訪者為 freshman」的條件下,計算「學生選擇步行 (Walk) 上學的機率」。
這是一個條件機率問題,表示為 。
首先,我們需要確定條件(分母)的樣本空間。題目指定「Given the respondents who were freshman」,所以我們的樣本空間是 freshman 的總人數。
從表格中:
freshman 搭乘捷運 = 70
freshman 搭乘公車 = 20
freshman 搭乘步行 = 65
freshman 的總人數 = 70 + 20 + 65 = 155
接下來,我們需要確定在 freshman 中,選擇步行上學的人數(分子)。
freshman 搭乘步行 = 65
第 11 題
- What of the following property is not appropriate to describe event A and event B are independent ?
A.
B.
C.
D.
E.
登入後即可作答並保存紀錄。
核心觀念
本題考「事件獨立」的定義及其等價條件。事件 與事件 獨立,定義為
若 、,亦可表示為
此外,由聯集公式:
代入獨立條件後得
題目中的 、 應按選項脈絡理解為條件機率 、;若原題確實省略條件機率符號,應視為排版簡寫。
解題方法
逐一比對各選項是否能描述獨立事件:
- 條件機率不因另一事件發生而改變,是獨立的等價條件。
- 交集機率等於機率乘積,是獨立的正式定義。
- 聯集公式代入獨立條件後,可得到選項 E。
- 表示兩事件互斥,通常與獨立不同。
因此,不適合描述獨立事件的是選項 C。
選項分析
A.
此式表示在事件 發生的條件下,事件 的機率沒有改變,正是 與 獨立的等價條件。
因此 A 正確。
B.
此式表示在事件 發生的條件下,事件 的機率沒有改變,也是獨立事件的等價條件。
因此 B 正確。
第 12 題
- A doctor is concerned about the relationship between blood pressure and sugar level. Among her patients, she classifies blood pressures as high, normal, or low and sugar level as normal and abnormal. She finds that (a) 15% have high blood pressure; (b) 20% have low blood pressure; (c) 18% have an abnormal sugar level; (d) of those with an abnormal sugar level, 35% have high blood pressure; and (e) of those with normal blood pressure, 12% have an abnormal sugar level. What percentage of her patients have a normal sugar level and low blood pressure?
A. 0.161
B. 0.176
C. 0.184
D. 0.920
E. 0.039
登入後即可作答並保存紀錄。
這題考驗對條件機率和聯合機率的計算,涉及多個事件的交集。我們可以使用表格法來整理資訊。
首先,定義事件:
- BP_H: High blood pressure (高血壓)
- BP_N: Normal blood pressure (正常血壓)
- BP_L: Low blood pressure (低血壓)
- SUGAR_A: Abnormal sugar level (異常血糖)
- SUGAR_N: Normal sugar level (正常血糖)
已知資訊:
(a)
(b)
(c)
(d)
(e)
我們需要計算 。
由 (a) 和 (b),我們可以推斷出 :
由於血壓只有高、正常、低三種情況,且是互斥的,總機率為 1。
由 (c),我們可以推斷出 :
血糖只有異常和正常兩種情況。
現在,利用條件機率公式 ,我們可以計算出一些聯合機率:
從 (d) :
從 (e) :
我們可以建立一個聯合機率表格:
| SUGAR_A (0.18) | SUGAR_N (0.82) | Total | |
|---|---|---|---|
| BP_H (0.15) | 0.063 | ? | 0.15 |
| BP_N (0.65) | 0.078 | ? | 0.65 |
| BP_L (0.20) | ? | ? | 0.20 |
| --------------- | ---------------- | ---------------- | ---------- |
| Total | 0.18 | 0.82 | 1.00 |
第 13 題
- (Continued Question 12.) Given patients with high blood pressure, what percentage of her patients have an abnormal sugar level?
A. 0.063
B. 0.12
C. 0.18
D. 0.35
E. 0.42
登入後即可作答並保存紀錄。
這題是上一題的延續,考驗條件機率的計算。
題目要求計算「在已知病人有高血壓 (BP_H) 的條件下,他們有異常血糖 (SUGAR_A) 的機率」,即 。
我們可以使用之前計算的聯合機率和邊際機率。
從上一題的表格或計算得知:
(已知條件 (a))
(從 (d) 計算得出)
條件機率的公式是:
代入數值:
進行計算:
(分子分母同乘以 1000,然後約分)
第 14 題
- The probability distribution of random variable, X, is defined as follows:
| X | Probability |
|---|---|
| 0 | 0.3 |
| 1 | ? |
| 2 | 0.1 |
| 3 | 0.3 |
| 4 | 0.25 |
What is the missing value?
A. 0
B. 0.05
C. 0.1
D. 0.15
E. 0.2
登入後即可作答並保存紀錄。
這題考驗對離散型隨機變數機率分佈基本性質的理解。
對於任何一個離散型隨機變數的機率分佈,其所有可能值的機率之和必須等於 1。
即:。
在給定的機率分佈中,我們有:
(未知值)
第 15 題
- Owing to the extravagant housing prices, the government would like to estimate the rate of the vacant house in the Taipei metropolitan area. From the past records, the rate of the vacant house is 10%. The government randomly selects 400 houses and assigns an agent to visit the house. Let X denote the number of non- vacant houses. What of the following condition is not necessary for X to be a binomial random variable?
A. A selected house is considered as a trial.
B. Each trial has only two outcomes such as "vacant" and "non-vacant".
C. Each trial has the same probability of the vacant house.
D. Trials are uncorrelated.
E. Trials are independent.
登入後即可作答並保存紀錄。
核心觀念
本題考查二項分配的成立條件。若隨機變數 服從二項分配,必須具備:
- 試驗次數固定為 。
- 每次試驗只有兩種結果。
- 每次試驗成功的機率相同,均為 。
- 各次試驗彼此獨立。
本題中令「選出的房屋為非空屋」視為成功,則
因此在二項分配模型下,
其機率質量函數為
解題方法
逐一比對各選項與二項分配的必要條件。關鍵在於「獨立」與「不相關」不是同一層次的條件:
- 獨立會推出不相關。
- 不相關不等於獨立。
- 二項分配要求的是各次試驗獨立,不要求額外列出「不相關」。
因此,判斷哪一項不是必要條件,即可找出「不相關」這一項。
選項分析
A. A selected house is considered as a trial.
正確。每選取一戶房屋,即進行一次試驗;共選取 戶,因此試驗次數固定為 次。這符合二項分配的固定試驗次數設定。
B. Each trial has only two outcomes such as "vacant" and "non-vacant".
正確。每戶房屋只有兩種分類結果:
- 空屋(vacant)
- 非空屋(non-vacant)
第 16 題
- (Continued Question 15.) Assume X is a binomial random variable. What of the following is a correct statement?
A. The mean number of vacant houses equals 40.
B. The mean number of non-vacant houses equals 15.
C. The probability that all of the house are non-vacant is .
D. The standard deviation of X equals 36.
E. The variance of X equals 6.
登入後即可作答並保存紀錄。
這題是上一題的延續,假設 X 是一個二項隨機變數,並要求找出正確的敘述。
從上一題的題目描述:
- 總房屋數 。
- 空屋率 (vacant rate) 。
- 定義為「非空屋 (non-vacant)」的數量。
因此,對於 X 這個二項隨機變數:
- 試驗次數 。
- 「成功」的定義是房屋是「非空屋」。
- 成功機率 。
- 失敗的機率 。
- 所以,。
現在我們來計算二項分佈的相關統計量:
- 平均數 (Mean): 。
- 變異數 (Variance): 。
- 標準差 (Standard Deviation): 。
現在我們來檢驗各個選項:
-
A. The mean number of vacant houses equals 40.
題目定義 是「非空屋」的數量。
「空屋」的數量是 。
空屋的平均數 。
所以,這個敘述是正確的。 -
B. The mean number of non-vacant houses equals 15.
我們計算的非空屋平均數是 。
15 與 360 相差甚遠。
所以,這個敘述是錯誤的。 -
C. The probability that all of the house are non-vacant is .
「所有房子都是非空屋」意味著 。
對於二項分佈 。
第 17 題
- Regarding a normal random variable X, what of the following statement is incorrect?
A. The mean and median of X are equal.
B. The mean and mode of X are equal.
C. The normal curve is symmetry about mode.
D. The probability below the mean equals 1/2.
E. The probability that X exceeds 1.96 equals 0.025.
登入後即可作答並保存紀錄。
這題考驗對常態分佈 (normal distribution) 性質的理解。題目要求找出錯誤的敘述。
常態分佈 的主要性質包括:
- 對稱性 (Symmetry): 分佈曲線關於平均數 對稱。
- 集中趨勢 (Central Tendency): 平均數 ()、中位數 (median) 和眾數 (mode) 都相等,且都位於對稱軸上。
- 形狀 (Shape): 呈鐘形曲線 (bell-shaped)。
- 機率性質 (Probability Properties):
- 平均數將總機率分成兩半:。
- 標準差 決定了曲線的寬度和分散程度。
- 從而, 且 。
現在我們逐一檢驗選項:
-
A. The mean and median of X are equal.
對於常態分佈,平均數、中位數和眾數都相等。此敘述為正確。 -
B. The mean and mode of X are equal.
同上,平均數和眾數也相等。此敘述為正確。
第 18 題
- The "fill" problem is important in many industries, such as those making cereal, toothpaste, beer, and so on. If an industry claims that it is selling 20 ounces of its product in a container, it must have a mean greater than 20 ounces, or else the FDA will crack down, although the FDA will allow a very small percentage of the containers to have less than 20 ounces. Assume the content X of a container follows a normal distribution with mean and variance . What is ?
A. 0.023
B. 0.484
C. 0.5
D. 0.516
E. 0.977
登入後即可作答並保存紀錄。
核心觀念
本題考查常態分配的標準化。若
則可透過
將 轉換為標準常態變數,再利用標準常態累積分配函數 計算機率。
題目給定:
因此標準差為
解題方法
要求的是
先標準化:
計算標準化後的 值:
所以
由標準常態分配表:
利用標準常態分配的對稱性:
因此
選項分析
第 19 題
- Which of the following is not a necessary condition for the well-known Central Limit Theorem?
A. All observations in the sample are independent.
B. Each observation in the sample has equal mean.
C. Each observation in the sample has equal variance.
D. Each observation in the sample has to have a normal distribution.
E. The sample size of the sample has to be large.
登入後即可作答並保存紀錄。
這題考驗對中央極限定理 (Central Limit Theorem, CLT) 的理解。
中央極限定理的核心內容是:
從任何分佈(無論是否為常態分佈)中抽取的一個足夠大的隨機樣本,其樣本平均數的分佈將近似於常態分佈。
CLT 的必要條件通常包括:
- 隨機抽樣 (Random Sampling): 樣本必須是隨機抽取的,這保證了每個樣本點都是從母體中隨機選取的。
- 獨立性 (Independence): 樣本中的各個觀測值 (observations) 必須是相互獨立的。這意味著一個觀測值的取值不影響其他觀測值的取值。
- 有限的平均數和變異數 (Finite Mean and Variance): 母體必須存在一個有限的平均數 和一個有限的變異數 。
- 足夠大的樣本數 (Sufficiently Large Sample Size): 樣本數 必須足夠大。通常 被認為是足夠大的,但具體大小取決於母體分佈的形狀。
現在我們來分析選項:
-
A. All observations in the sample are independent. (樣本中的所有觀測值都是獨立的。)
這是 CLT 的一個關鍵要求。獨立性保證了樣本點之間沒有系統性的關聯,這對於樣本平均數的期望值和變異數的計算很重要。這是必要的。 -
B. Each observation in the sample has equal mean. (樣本中的每個觀測值都有相同的平均數。)
這句話的表述有點奇怪。「每個觀測值」本身是一個隨機變數。如果指每個觀測值來自同一個母體,那麼它們的期望值(平均數)應該是相同的,也就是母體的平均數 。這實際上是有限母體平均數的性質。如果指樣本平均數的平均數,那也是 。這個條件與母體平均數有限的要求有關,但表述方式不夠清晰。
第 20 題
- Which of the following is not a necessary condition for a random sample?
A. All observations are correlated.
B. All observations are drawn from the same population.
C. Each observation has equal mean.
D. Each observation has equal variance.
E. Each observation has a common distribution.
登入後即可作答並保存紀錄。
這題考驗對「隨機樣本」(random sample) 的定義和條件的理解。
隨機樣本是指從一個母體中抽取的一組樣本,其中每個可能的樣本都有相同的機率被抽取。更具體地,對於簡單隨機樣本 (simple random sample),通常要求:
- 來自同一母體 (Drawn from the same population): 樣本中的所有觀測值都應該是從同一個母體中抽取出來的。
- 獨立性 (Independence): 樣本中的觀測值之間應該是相互獨立的(或至少是不相關的)。
我們來分析各選項:
-
A. All observations are correlated. (所有觀測值都是相關的。)
隨機樣本的一個關鍵特徵是,樣本點之間通常應該是獨立的(或至少是不相關的)。「相關」(correlated) 表示它們之間存在某種線性關係。如果所有觀測值都是相關的,這就違反了隨機樣本的獨立性要求。因此,這不是隨機樣本的必要條件,反而是一個不符合隨機樣本特性的描述。所以,這個敘述本身是錯誤的,因此可以說它「不是隨機樣本的必要條件」,因為隨機樣本要求的是獨立或不相關。 -
B. All observations are drawn from the same population. (所有觀測值都來自同一個母體。)
這是隨機樣本的定義之一。樣本必須代表其所來自的母體。這是必要的。 -
C. Each observation has equal mean. (每個觀測值都有相同的平均數。)
如果樣本是從同一個母體中抽取出來的,那麼每個觀測值(作為一個隨機變數)的期望值都應該等於母體的平均數 。
第 21 題
- What of the following is not the property for the sample mean?
A. The mean of the sample mean equals the population mean.
B. The observations used to derive the sample mean have to be uncorrelated.
C. The population mean can be estimated by it.
D. The standard deviation of the sample mean is a function of the sample size.
E. When the sample size is large, the distribution of the sample mean is approximately normal.
登入後即可作答並保存紀錄。
核心觀念
本題考查「樣本平均數(sample mean)」的基本性質。若 為樣本,樣本平均數定義為
在隨機抽樣且各觀測值獨立同分布(i.i.d.)的常見條件下,若母體平均數為 、母體變異數為 ,則:
-
期望值
-
變異數與標準差
因此樣本平均數的標準差為
-
大樣本分布
依據中央極限定理,當樣本數 夠大時,樣本平均數的分布近似常態分布:
本題要求找出「不是樣本平均數的性質」者。
解題方法
逐一檢查各選項是否為樣本平均數的基本性質。判斷重點在於:
- 樣本平均數是否為母體平均數的無偏估計量;
- 樣本平均數的變異數是否受到樣本數影響;
- 大樣本下是否具有近似常態性;
- 觀測值「必須不相關」是否真的是樣本平均數成立的必要條件。
其中,樣本平均數本身只是觀測值的算術平均:
觀測值是否不相關,並不是「能夠計算或定義樣本平均數」的必要條件。因此,B 的敘述過於絕對。
選項分析
A. The mean of the sample mean equals the population mean.
此敘述正確。
由期望值的線性性質:
若每個觀測值的期望值皆為母體平均數 ,則:
因此,樣本平均數是母體平均數 的無偏估計量。
B. The observations used to derive the sample mean have to be uncorrelated.
此敘述錯誤,為本題答案。
樣本平均數的定義為:
不論觀測值之間是否相關,都可以計算樣本平均數。因此,「觀測值必須不相關」不是樣本平均數的必要性質。
若觀測值彼此相關,樣本平均數的變異數會包含共變異數項:
第 22 題
- Which of the following is a correct statement when the sample size increases?
A. The length of the confidence interval of the population mean increases.
B. The population mean increases.
C. The sample mean increases.
D. The standard deviation of the population decreases.
E. The standard deviation of the sample mean decreases.
登入後即可作答並保存紀錄。
這題考驗對信賴區間 (confidence interval) 和樣本平均數標準誤 (standard error of sample mean) 的理解,以及樣本數 (sample size) 對它們的影響。
當樣本數 增加時:
-
A. The length of the confidence interval of the population mean increases. (母體平均數信賴區間的長度增加。)
信賴區間的長度通常與樣本數 的關係是反比的。
例如,對於平均數的 信賴區間: (母體變異數已知) 或 (母體變異數未知)。
區間長度為 或 。
由於 在分母,當 增加時, 減少,因此信賴區間的長度會減少。
所以,這個敘述是錯誤的。 -
B. The population mean increases. (母體平均數增加。)
母體平均數 是母體的固定參數,它是一個常數,不隨樣本數 的變化而變化。
所以,這個敘述是錯誤的。 -
C. The sample mean increases. (樣本平均數增加。)
樣本平均數 是從樣本計算出來的統計量。它是一個隨機變數。
第 23 題
- The Tourism Bureau of Taiwan would like to estimate the mean amount of money spent by a tourist to within 2500, what sample size would be necessary for the Chamber of Commerce to meet its objective in estimating the mean spending amount?
A. 5
B. 7
C. 17
D. 25
E. 44
登入後即可作答並保存紀錄。
這題考驗樣本數的計算 (sample size calculation),用於估計母體平均數。
題目要求估計母體平均數 ,並給定了以下資訊:
- 估計精度 (Margin of Error, E): 。這意味著我們希望 。
- 信賴水準 (Confidence Level): 90%。
- 母體標準差 (Population Standard Deviation, ): 。
- 母體分佈: 假設為常態分佈 (normally distributed)。
對於常態分佈的母體,估計母體平均數所需的樣本數 可以通過以下公式計算:
其中:
- 是允許的最大誤差 (margin of error)。
- 是對應於所需信賴水準的標準常態分佈的臨界值。
- 是母體標準差。
- 是樣本數。
首先,我們需要找到 。
信賴水準是 90%,所以 。
。
我們需要查找 ,即標準常態分佈中,右尾機率為 0.05 的 z 值。
從標準常態分佈表中,與機率 0.95 (即 ) 對應的 z 值約為 1.645。
所以,。
第 24 題
- What of the following is a correct property for the confidence interval?
A. The length of the confidence interval is not related to the size of the random sample.
B. The length of confidence interval is shorter as the mean gets larger.
C. The length of confidence interval is shorter as the sample size increases.
D. The length of confidence interval is shorter as the standard deviation gets larger.
E. The length of confidence interval is shorter when the confidence level gets larger.
登入後即可作答並保存紀錄。
這題考驗對信賴區間 (confidence interval) 長度性質的理解。
信賴區間的長度通常由以下公式決定(以母體平均數為例,母體變異數已知):
Length =
其中:
- 與信賴水準 有關。
- 是母體標準差。
- 是樣本數。
我們來分析各選項:
-
A. The length of the confidence interval is not related to the size of the random sample. (信賴區間的長度與隨機樣本的大小無關。)
從公式 Length = 可知,長度與 成正比。因此,長度與樣本數 顯然有關。此敘述是錯誤的。 -
B. The length of confidence interval is shorter as the mean gets larger. (信賴區間的長度隨著平均數的增大而變短。)
信賴區間的長度與母體平均數 (或樣本平均數 ) 無關。它與 , , 和信賴水準有關。此敘述是錯誤的。 -
C. The length of confidence interval is shorter as the sample size increases. (隨著樣本數的增加,信賴區間的長度變短。)
第 25 題
- Which is the least appropriate statistical test to evaluate the association between the data for the level of education for the head of household measured as less than high school (HS) and greater than HS and household income (unit: thousand dollars)?
A. Chi-square test
B. Mann-Whitney test
C. One-way ANOVA
D. Regression
E. Two-independent sample T test
登入後即可作答並保存紀錄。
這題考驗對統計檢定方法適用性的判斷,特別是當變數類型不同時。
題目要評估的變數是:
- 教育程度 (Level of education): 這是二分類變數(少於高中 vs. 高於高中)。可以視為名目尺度或次序尺度。
- 家庭收入 (Household income): 這是數值型變數(比例尺度)。
我們需要尋找一個最不恰當的統計檢定方法來評估這兩個變數之間的關聯性。
-
A. Chi-square test (卡方檢定): 卡方檢定主要用於檢定兩個類別型變數之間是否存在關聯性,或者檢定觀察頻率與期望頻率之間是否存在差異。當一個變數是類別型(如教育程度)而另一個是數值型時,我們通常會先將數值型變數分組(離散化),使其變成類別型,然後再進行卡方檢定。或者,更常見的是將類別型變數轉換成虛擬變數 (dummy variables) 來與數值型變數進行分析。雖然不是最直接的方法,但卡方檢定可以應用。
-
B. Mann-Whitney test (曼-惠特尼 U 檢定): 這是一種非參數檢定,用於比較兩個獨立樣本的中位數是否存在顯著差異。它適用於次序尺度或連續尺度資料。在這裡,我們可以將教育程度視為分組變數(兩組:少於高中 vs. 高於高中),然後檢定這兩組在家庭收入(數值型)的中位數上是否有差異。這是適用的。
-
C. One-way ANOVA (單因子變異數分析): ANOVA 用於比較三個或更多組的平均數是否存在顯著差異。在這裡,教育程度是二分類變數,可以視為一個因子,將收入分成兩組進行比較。這相當於獨立樣本 t 檢定(當組別為兩組時)。因此,ANOVA 是適用的。
-
D. Regression (迴歸分析): 迴歸分析可以用來模型化一個依變數(如家庭收入)與一個或多個自變數(如教育程度)之間的關係。我們可以將教育程度編碼(例如:0 代表少於高中,1 代表高於高中),然後建立一個線性迴歸模型:Income = 。
第 26 題
- Which is the most appropriate statistical model or test to evaluate the association between the data for the floor area (unit: square foot) and household income (unit: thousand dollars)?
A. Chi-square test
B. Mann-Whitney test
C. One-way ANOVA
D. Regression
E. Two-independent sample T test
登入後即可作答並保存紀錄。
這題考驗對兩個數值型變數 (numerical variables) 之間關聯性評估的統計模型選擇。
題目要評估的變數是:
- 樓地板面積 (Floor area): 數值型變數(比例尺度)。
- 家庭收入 (Household income): 數值型變數(比例尺度)。
我們需要尋找最恰當的統計模型或檢定方法來評估這兩個數值型變數之間的關聯性。
-
A. Chi-square test (卡方檢定): 通常用於類別型變數之間的關聯性檢定。對於兩個數值型變數,需要先將它們離散化成類別,這樣會損失資訊,所以不是最恰當的方法。
-
B. Mann-Whitney test (曼-惠特尼 U 檢定): 用於比較兩個獨立樣本的中位數。它需要一個分組變數(類別型)和一個數值變數。不適用於兩個數值變數之間的關聯性。
-
C. One-way ANOVA (單因子變異數分析): 用於比較一個類別型變數(因子)對一個數值型變數的影響。不適用於兩個數值型變數之間的關聯性。
第 27 題
- The average time for a resident in Taiwan watches TV in Taiwan about 2 hours per day. The following provides some summary statistics for a sample collected from a survey in Japan:
| Respondents | N | Sample mean | Sample standard deviation |
|---|---|---|---|
| 100 | 2.5 | 0.6 |
95% CI for average: (2.38, 2.62)
Which of the following statement is the most appropriate statement?
A. The average time for a resident in Taiwan watching TV is significantly different than that for past research.
B. The average time for a resident in Taiwan watching TV is significantly higher than that for past research.
C. The average time for a resident in Taiwan watching TV is significantly lower than that for past research.
D. The average time for a resident in Taiwan watching TV is not significantly different from that for the past research.
E. The average time for a resident in Taiwan watching TV exactly equals to 2.5.
登入後即可作答並保存紀錄。
核心觀念
本題考查:
- 母體平均數的信賴區間。
- 以信賴區間判斷假設檢定結果。
- 統計顯著性與差異方向的判斷。
研究過去指出,每日平均看電視時間為 小時,因此設定:
其中 表示居民每日平均看電視時間。
題目提供樣本的 信賴區間為:
若假設的母體平均數 不在 信賴區間內,則在顯著水準 下,可拒絕 ,判定母體平均數與 小時有顯著差異。
解題方法
樣本統計量為:
樣本平均數的標準誤為:
信賴區間可表示為:
代入數值:
因此:
過去研究的平均值為 小時,而:
所以 不在信賴區間 內,代表樣本結果顯示母體平均看電視時間與過去研究的 小時有顯著差異。
此外,整個信賴區間都大於 ,且樣本平均數 ,因此差異方向是「顯著較高」。
故最適當的敘述為 B。
選項分析
A. The average time for a resident in Taiwan watching TV is significantly different than that for past research.
此敘述指出平均看電視時間與過去研究有顯著差異,方向沒有說明。由於 不在 信賴區間內,此敘述的「有顯著差異」本身成立。
但本題信賴區間完全高於 ,可以進一步判斷是「顯著較高」,因此 B 比 A 更精確、也更適當。選擇題要求最適當敘述時,不選 A。
第 28 題
- Suppose the city government would like to build a new metro route. They would like to know whether the local resident approves this policy. A survey is distributed to understand the residents' opinion. Let the probability of approving this policy be denoted as p. What is the least sample size that the city government needs to have the marginal error of E=0.06 with 95% confidence?
A. 67
B. 188
C. 267
D. 534
E. 376
登入後即可作答並保存紀錄。
核心觀念
本題考查母體比例 的信賴區間與樣本數估計。
當以樣本比例 估計母體比例 時,在大樣本下有
其中:
- :邊際誤差(margin of error)
- :標準常態分配的臨界值
- :樣本數
因此,樣本數公式為
題目未提供 的先驗估計值時,取 ,因為
且在 時達到最大值,可確保樣本數足以符合誤差要求。
解題方法
題目要求:
- 信賴水準為
- 邊際誤差為
- 未知母體比例
信賴水準對應的標準常態臨界值為
未知 時採用最保守估計 ,故 。
代入樣本數公式:
第 29 題
- Let and denote the mean of the house price for urban and suburban areas. Also, let and denote the sample mean of the house price for suburban and rural areas. Suppose a potential customer would like to evaluate whether the house price (unit: thousand dollars) in the rural area is much lower than that in the suburban. Which of the following statement is correct?
A. versus .
B. versus .
C. versus .
D. versus .
E. versus .
登入後即可作答並保存紀錄。
核心觀念
本題考「單一方向的兩母體平均數假設檢定」。
研究問題是: rural area 的房價是否低於 suburban area。因此應將研究假設設定為:
虛無假設則包含「沒有較低」的所有情況:
假設檢定的基本原則是:
- 研究者想證明的主張放在對立假設 。
- 虛無假設 通常包含等號。
- 檢定母體平均數時,應使用 ,而非樣本平均數 。
題目中的符號敘述有筆誤:前文對 的區域定義,與 的區域定義不一致。依選項設計與題意,應將 視為 rural area 的母體平均房價, 視為 suburban area 的母體平均房價。
解題方法
題目要判斷 rural 房價是否「低於」suburban 房價,因此研究方向是左尾檢定:
其中:
- :rural area 的母體平均房價
- :suburban area 的母體平均房價
對立假設為:
其互補事件即為虛無假設:
所以正確的假設設定為:
選項分析
A. versus
錯誤。 與 完全相同,兩個假設沒有互補性,也無法形成有效的假設檢定。
此外,本題是單尾檢定,不是只檢定兩者是否完全相等。
第 30 題
- Which of the following is a correct statement about the null hypothesis that investigate the association between the method of commute and gender?
A. : the commuting method and gender are dependent.
B. : the commuting method and gender are independent.
C. : the commuting method and gender are not significantly associated.
D. : the commuting method and gender are significantly associated.
E. : the commuting method and gender are uncorrelated.
登入後即可作答並保存紀錄。
核心觀念
本題考查「類別變數獨立性檢定」的虛無假設。
通勤方式與性別皆屬於類別變數,因此可使用卡方獨立性檢定,檢驗兩者在母體中是否具有關聯:
- 虛無假設 :通勤方式與性別相互獨立。
- 對立假設 :通勤方式與性別不獨立,亦即具有關聯。
若觀察次數與在獨立假設下的期望次數差異過大,便拒絕 ,認為通勤方式與性別存在顯著關聯。
卡方統計量為
其中 為觀察次數, 為獨立假設下的期望次數:
解題方法
題目詢問「檢驗通勤方式與性別關聯性時,虛無假設為何」。
獨立性檢定的標準設定是:
因此,直接尋找表示「兩個類別變數獨立」的選項即可。
選項分析
A. : the commuting method and gender are dependent.
錯誤。Dependent 表示「相依」,代表通勤方式與性別存在關聯,應屬於對立假設 ,不是虛無假設。
第 31 題
- Which is the most appropriate test to evaluate the association between the commuting method (nominal scale) and gender?
A. Chi-square test
B. F test
C. Mann-Whitney test
D. Paired T test
E. Two-independent sample T test
登入後即可作答並保存紀錄。
這題考驗對適合評估兩個類別型變數 (categorical variables) 之間關聯性的統計檢定方法的選擇。
題目要評估的變數是:
- 通勤方式 (Commuting method): 名目尺度 (nominal scale) 的類別型變數。
- 性別 (Gender): 名目尺度 (nominal scale) 的類別型變數。
我們需要尋找最恰當的統計檢定方法來評估這兩個類別型變數之間的關聯性。
-
A. Chi-square test (卡方檢定): 卡方檢定,特別是卡方獨立性檢定 ( test of independence),是專門用來檢定兩個類別型變數之間是否存在關聯性的標準統計方法。它通過比較觀察到的頻率分佈與在獨立假設下期望的頻率分佈來進行。這是最恰當的方法。
-
B. F test (F 檢定): F 檢定通常用於變異數分析 (ANOVA) 或檢定迴歸模型的顯著性,涉及數值型變數。它不直接用於檢定兩個類別型變數之間的關聯性。
第 32 題
- Let denote the population mean of IQ. Which of the following is a correct statement?
A. For a fixed significance level , the probability of a Type I error decreases when the sample size increases.
B. If we reject in a study about change in weight on a new diet using a significance level of , then we also reject it using .
C. In the test of versus , the p-value is greater than greater than 0.01. Then, a 99% confidence interval for should be (120, 130).
D. .
E. The p value is defined as the probability that is true.
登入後即可作答並保存紀錄。
這題考驗對假設檢定 (hypothesis testing) 中各種概念的理解,包括第一類錯誤、第二類錯誤、p-value、信賴區間和樣本數的影響。
-
A. For a fixed significance level , the probability of a Type I error decreases when the sample size increases. (在固定的顯著水準 下,犯第一類錯誤的機率隨著樣本數的增加而減少。)
第一類錯誤 (Type I error) 是指當虛無假說 () 為真時,拒絕 的機率。這個機率就是顯著水準 。
。
的設定是獨立於樣本數 的。無論樣本數是多少,只要 固定,犯第一類錯誤的機率就固定為 。
樣本數的增加主要影響檢定的「統計檢定力」(power),即正確拒絕錯誤的 的機率 (),或者減少犯第二類錯誤 () 的機率。
所以,這個敘述是錯誤的。 -
B. If we reject in a study about change in weight on a new diet using a significance level of , then we also reject it using . (如果在 的顯著水準下拒絕 ,那麼在 的顯著水準下也會拒絕它。)
這裡的 寫法不標準,通常是 和 。假設這裡的 是指 (或 的反面)。
如果我們在 下拒絕 ,這意味著 p-value 。
如果 p-value ,那麼它也必然 ,因為 。
所以,在 下,我們也會拒絕 。
這個敘述是正確的。 -
C. In the test of versus , the p-value is greater than greater than 0.01. Then, a 99% confidence interval for should be (120, 130). (在檢定 對 時,p-value 大於 0.01。那麼, 的 99% 信賴區間應為 (120, 130)。)
這裡有兩個關鍵點:- p-value > 0.01
- 99% CI for is (120, 130)。
虛無假說 。
第 33 題
- An industrial plant claims to discharge no more than 1000 gallons of wastewater per hour, on the average, into a neighboring lake. An environmental action group took some random samples from the lake over a period of time to analyze the true mean discharge of wastewater per hour. Consider a one-sided hypotheses and use a significance level of . Which of the following statement is correct?
A. If is rejected, the probability that is actually true is 0.05.
B. If was true, the probability that the sample mean equals exactly the observed value would equal 0.05.
C. If the plant exceeds the limit, there is only a 5% chance that we will conclude that they do not exceed the limit.
D. If the plant does not exceed the limit, there is only 5% chance that we will conclude they exceed the limit.
E. We have 5% confidence that this plant does not discharge more than 1000 gallons of wastewater.
登入後即可作答並保存紀錄。
核心觀念
本題考查單尾假設檢定中的「顯著水準 」與第一類錯誤。
令母體平均每小時排放量為 ,則假設可設為:
因為環保團體要判斷工廠是否超過每小時 加侖的限制,所以拒絕域位於右尾。
顯著水準定義為:
本題 ,表示在工廠實際沒有超過排放上限時,卻誤判為超過上限的機率為 。這種錯誤稱為第一類錯誤。
解題方法
判斷每個選項是否符合顯著水準的條件定義:
在本題中:
- 「拒絕 」代表判定工廠超過每小時 加侖。
- 「 為真」代表工廠實際排放量沒有超過每小時 加侖。
因此:
這正是選項 D 的敘述。
選項分析
A. If is rejected, the probability that is actually true is 0.05.
錯誤。
此敘述將顯著水準誤解為「在拒絕 的條件下, 為真的機率」:
但顯著水準所表示的是:
兩者條件方向不同,不能直接互換。選項 A 是典型的條件機率倒置錯誤。
B. If was true, the probability that the sample mean equals exactly the observed value would equal 0.05.
錯誤。
顯著水準不是指樣本平均數恰好等於某個觀察值的機率。若樣本平均數服從連續型分配,則對任何特定數值 ,通常有:
假設檢定關注的是樣本統計量落入拒絕域的機率,例如:
而不是樣本平均數「恰好等於觀察值」的機率。
第 34 題
- Let and denote the probability of possessing a house for female and male in Taiwan. Suppose that a 95% confidence interval for the difference equals (-0.05, -0.01). What of the following statement is correct?
A. The probability of possessing a house for male and female is not significantly different.
B. The difference in the probability of possessing a house between male and female is between -0.05 and -0.01.
C. We are 95% confident that the probability of possessing a house for female is significantly larger than that for male.
D. We are 95% confident that a minority of Taiwan residents possesses a house.
E. We can conclude that the population proportions for male and female are equal.
登入後即可作答並保存紀錄。
這題考驗對比例差異信賴區間 (confidence interval for difference in proportions) 的解釋。
題目設定:
- : 女性擁有房子的機率 (population proportion)。
- : 男性擁有房子的機率 (population proportion)。
- 95% 信賴區間 (CI) 為 的差值:(-0.05, -0.01)。
信賴區間 (-0.05, -0.01) 表示,我們有 95% 的信心認為,男性擁有房子的機率與女性擁有房子的機率之差 () 落在 -0.05 到 -0.01 之間。
現在我們來分析各選項:
-
A. The probability of possessing a house for male and female is not significantly different. (男性和女性擁有房子的機率沒有顯著差異。)
顯著性差異的判斷通常與信賴區間是否包含 0 有關。
在這裡,我們計算的是 的信賴區間。
由於信賴區間 (-0.05, -0.01) 完全不包含 0,且整個區間都是負值,這表明 顯著地小於 0。
換句話說,男性擁有房子的機率顯著低於女性擁有房子的機率。
所以,它們之間存在顯著差異。此敘述是錯誤的。 -
B. The difference in the probability of possessing a house between male and female is between -0.05 and -0.01. (男性和女性之間擁有房子的機率差異在 -0.05 和 -0.01 之間。)
信賴區間 (-0.05, -0.01) 就是對母體參數(在這裡是 )的估計範圍。
這個敘述直接描述了信賴區間的含義。
此敘述是正確的。 -
C. We are 95% confident that the probability of possessing a house for female is significantly larger than that for male. (我們有 95% 的信心認為,女性擁有房子的機率顯著大於男性。)
信賴區間是 ,表示 是負值。
第 35 題
- Let denote the population mean for the ith group, . When the one-way ANOVA provides relatively more evidence that is rejected, which of the following statement is appropriate?
A. The between-groups variation is smaller.
B. The between-groups variation is larger.
C. The within-groups variation is smaller.
D. The within-groups variation is larger.
E. The total sum of square is unchanged.
登入後即可作答並保存紀錄。
這題考驗對單因子變異數分析 (One-way ANOVA) 結果的解釋。
ANOVA 的目的是檢定多個群體( 個群體)的母體平均數是否存在顯著差異。
虛無假說 () 是所有群體的母體平均數都相等:。
對立假說 () 是至少有一個群體的母體平均數不同。
當 ANOVA 檢定提供「相對較多證據」表明 被拒絕時,這意味著我們有足夠的統計證據來推斷至少有一個群體的平均數與其他群體不同。換句話說,群體之間的平均數存在顯著差異。
ANOVA 的檢定統計量 F 值是通過比較組間變異數 (between-groups variation) 和組內變異數 (within-groups variation) 來計算的:
- 組間變異數 (Between-groups variation, SSBetween 或 SSB): 衡量各組平均數與總平均數之間的差異。它反映了不同組別之間的變異。
- 組內變異數 (Within-groups variation, SSWithin 或 SSW): 衡量每個觀測值與其所在組平均數之間的差異。它反映了組內的隨機變異或誤差。
當 被拒絕時,意味著組間變異數相對於組內變異數是比較大的。
也就是說,MSB 相對於 MSW 較大,導致 F 值較大。
因此,如果 被拒絕,表示群體平均數之間存在顯著差異,這主要是由較大的組間變異所致。
- A. The between-groups variation is smaller. (組間變異數較小。)
如果組間變異數較小,F 值就會較小,不利於拒絕 。這是錯誤的。
第 1 題
一、True and False(是非題) questions. [State the reason for your answer. 3% for each,30%]
- A distribution table is required when using the p value of a test statistic to make a decision.
登入後即可作答並保存紀錄。
這題考驗對假設檢定中 p-value 的理解。
p-value 的定義是在虛無假說為真的前提下,觀察到樣本結果或更極端結果的機率。它的計算通常是基於檢定統計量 (test statistic) 的值及其理論分佈。
雖然在傳統的統計方法中,我們可能會查閱標準常態分佈表、t 分佈表或卡方分佈表來尋找 p-value 所對應的機率值,但 p-value 本身是一個機率值,其計算可以透過統計軟體直接得到,不一定「必須」依賴於某種「分佈表」。分佈表是輔助工具,用於查找特定統計量值對應的機率,或特定機率值對應的統計量值。
第 2 題
- A random variable cannot be negative.
登入後即可作答並保存紀錄。
這題考驗對隨機變數 (random variable) 值域的理解。
隨機變數是將隨機實驗的結果映射到實數的函數。其值域 (range) 取決於隨機實驗的性質。
有些隨機變數的值域確實是有限制,不能為負,例如:
- 計數型隨機變數(如:某時間段內到達的顧客數),其值只能是非負整數 (0, 1, 2, ...)。
- 表示機率或比例的隨機變數,其值域在 [0, 1] 之間。
第 3 題
- Assuming a random variable X has a normal distribution with mean and standard deviation , then .
登入後即可作答並保存紀錄。
這題考驗對常態分佈 (normal distribution) 對稱性的理解。
常態分佈是一個以平均數 為中心的對稱分佈。這意味著分佈曲線在 處左右對稱。
對於任何對稱分佈,平均數(或中位數)將總機率 1 分成兩半。
因此:
- (落在平均數左側的機率)
- (落在平均數右側的機率)
- 對於連續型隨機變數,任何單一點的機率為 0,所以 。
第 4 題
- Consider two events A and B and suppose , then .
登入後即可作答並保存紀錄。
這題考驗對條件機率 (conditional probability) 的理解,以及事件之間的關係(關聯性、獨立性)。
表示在事件 B 已經發生的條件下,事件 A 發生的機率。
是事件 A 發生的機率。
表示事件 B 的發生增加了事件 A 發生的機率,即 A 和 B 之間存在正向關聯。
然而,這個敘述並非總是成立的。事件 A 和 B 之間可能存在以下關係:
- 正向關聯 (Positive association): 。
- 負向關聯 (Negative association): 。
- 獨立 (Independence): 。
題目中的敘述只描述了第一種情況,而忽略了其他兩種可能性。因此,該敘述不一定為真。
第 5 題
- Consider two events A and B, then we have .
登入後即可作答並保存紀錄。
這題考驗對機率加法法則 (addition rule of probability) 的理解。
機率的加法法則指出:
其中 是事件 A 和 B 同時發生的機率。
由於機率值永遠是非負的,即 ,
因此,我們可以從加法法則推導出不等式:
。
第 6 題
- Events A and B are mutually exclusive which imply they are independent.
登入後即可作答並保存紀錄。
這題考驗對「互斥事件」(mutually exclusive events) 和「獨立事件」(independent events) 的定義及其關係的理解。
- 互斥事件: 兩個事件 A 和 B 不能同時發生,即 。
- 獨立事件: 事件 A 的發生與否不影響事件 B 發生的機率,即 。
題目聲稱「互斥事件意味著它們是獨立的」。
如果 A 和 B 是互斥事件,則 。
如果它們同時也是獨立事件,則必須滿足 。
因此,必須有 。
第 7 題
- The probability of an event A has to be less than 1.
登入後即可作答並保存紀錄。
這題考驗對機率基本性質的理解。
機率的定義域要求任何事件的機率 都必須滿足:
其中:
- 表示事件 E 是不可能事件 (impossible event)。
- 表示事件 E 是必然事件 (certain event)。
第 8 題
- The shape of the binomial distribution is symmetry when the probability of success equals 0.5.
登入後即可作答並保存紀錄。
這題考驗對二項分佈 (binomial distribution) 形狀特徵的理解。
二項分佈 描述了在 次獨立的伯努利試驗中,成功次數 的機率。其中 是單次試驗成功的機率。
二項分佈的形狀取決於參數 和 :
- 當 時:
二項分佈是對稱的。因為成功和失敗的機率相等,所以分佈的機率質量函數 (PMF) 在 的兩側是對稱的,具體來說,。
第 9 題
- The significance level of a test statistic has to be given in advance when performing the hypothesis test problem.
登入後即可作答並保存紀錄。
這題考驗對假設檢定 (hypothesis testing) 中顯著水準 (significance level, ) 設定的理解。
顯著水準 代表了我們願意承擔犯第一類錯誤 (Type I error) 的最大機率,即當虛無假說 () 為真時,卻拒絕 的機率。
在進行假設檢定前,必須預先設定 的值(例如 0.05 或 0.01)。這樣做有幾個重要原因:
- 客觀性: 預先設定 確保了檢定的客觀性。如果在看到 p-value 後才決定 ,決策可能會受到結果的影響而變得主觀。
第 10 題
- When the correlation of coefficient of random variables X and Y equals zero, we have X and Y are independent.
登入後即可作答並保存紀錄。
這題考驗對相關係數 (correlation coefficient) 與獨立性 (independence) 之間關係的理解。
相關係數 ( 或 ) 衡量的是兩個隨機變數之間的線性關係強度和方向。
- 表示兩個變數之間沒有線性關係。
獨立性是一個更強的概念。如果兩個隨機變數 X 和 Y 是獨立的,則它們之間不存在任何形式的關係,包括線性關係和非線性關係。
關鍵點:
- 獨立 相關係數為 0: 如果 X 和 Y 是獨立的,它們之間不可能存在線性關係,所以相關係數必定為 0。
- 相關係數為 0 獨立: 如果 X 和 Y 的相關係數為 0,這只表示它們之間沒有線性關係。但它們之間可能存在非線性關係。
第 1 題
- What is the data type of the floor area (unit: in square foot)?
A. Count
B. Interval
C. Nominal
D. Ordinal
E. Ratio
登入後即可作答並保存紀錄。
這題考驗對資料類型 (data types) 的認識。
資料類型依據測量尺度分為:名目 (Nominal)、次序 (Ordinal)、區間 (Interval) 和比例 (Ratio)。
- 名目尺度: 僅用於分類,無順序。
- 次序尺度: 有順序,但等級間差距不確定。
- 區間尺度: 有順序,等級間差距固定,但無絕對零點。
- 比例尺度: 有順序,等級間差距固定,有絕對零點,可進行比例運算。
「樓地板面積」(floor area),單位是「平方英尺」,具有以下特徵:
- 絕對零點: 0 平方英尺表示沒有面積。
第 2 題
- What is the data type of the type of transportation to work such as walk, bus, metro or car?
A. Count
B. Interval
C. Nominal
D. Ordinal
E. Ratio
登入後即可作答並保存紀錄。
這題考驗對資料類型的認識,特別是分類資料。
「上班的交通方式」(步行、公車、捷運、汽車)是對類別的分類。
這些類別之間沒有內在的數學順序或等級關係。雖然我們可以人為地對它們進行排序(例如,按花費時間或金錢),但這種排序並非資料本身的屬性。
- 名目尺度 (Nominal): 資料用於分類,類別之間無順序。
- 次序尺度 (Ordinal): 資料有順序,但等級間差距不確定。
第 3 題
- What is the data type for the preference of a certain type of coffee measured in a five level likert scale (like very much, like, fair, dislike, very dislike)?
A. Count
B. Interval
C. Nominal
D. Ordinal
E. Ratio
登入後即可作答並保存紀錄。
這題考驗對李克特量表 (Likert scale) 資料類型的判斷。
李克特量表用於測量態度、意見或偏好,它提供一系列等級供受訪者選擇。
題目中的五點李克特量表(非常喜歡、喜歡、普通、不喜歡、非常不喜歡)具有以下特徵:
- 分類: 將偏好歸入五個等級。
- 順序: 等級之間有明確的順序。
- 等級差距: 等級之間的差距通常假設是相等的,但這不一定嚴格成立。
- 無絕對零點: 沒有「零偏好」或「無偏好」的絕對零點。
第 4 題
- Which of the following is the most appropriate statistic to summarize the type of transportation to work such as walk, bus, metro or car?
A. Interquartile
B. Mean
C. Median
D. Range
E. None of the above
登入後即可作答並保存紀錄。
核心觀念
「通勤方式」分為步行、公車、捷運、汽車等類別,屬於名目尺度資料。各類別只有名稱上的差異,沒有大小順序,也不能進行加減或計算平均數。
描述名目資料時,適合使用眾數,也就是出現次數最多的類別。題目選項中沒有眾數,因此應選「以上皆非」。
解題方法
先判斷資料的尺度,再確認統計量是否適用:
- 通勤方式是類別資料,沒有數值大小或先後次序。
- 平均數、中位數、全距與四分位距都需要數值資料,部分還需要資料具有順序。
- 眾數可用於名目資料,但不在選項中。
因此,選項中的統計量都不適合;最適當的答案是 E。
選項分析
- A. 四分位距(Interquartile):四分位距是第三四分位數減第一四分位數,用來描述數值資料中央一半的分散程度。
第 5 題
- Which of the following is the least appropriate statistical statistic to summary the central tendency of the data for the floor area (unit: square foot)?
A. Mean
B. Median
C. Midpoint
D. Midrange
E. Mode
登入後即可作答並保存紀錄。
這題考驗對集中趨勢 (central tendency) 統計量在比例尺度資料上的適用性。
題目問的是「樓地板面積」(比例尺度)的集中趨勢。
- A. Mean (平均數): 對於比例尺度資料,平均數是最常用且適當的集中趨勢指標。
- B. Median (中位數): 對於比例尺度資料,中位數也是一個有效的指標,尤其是在有離群值時。
- C. Midpoint (中點): 這個詞較模糊,若指區間中點,則非整體集中趨勢。
- D. Midrange (中距): (最小值+最大值)/2。
第 6 題
- Which of the following is the lease appropriate statistic to summary the variability of the data for the floor area (unit: square foot)?
A. Correlation
B. Interquartile
C. Range
D. Standard deviation
E. Variance
登入後即可作答並保存紀錄。
這題考驗對變異性 (variability) 統計量在比例尺度資料上的適用性。
題目問的是「樓地板面積」(比例尺度)的變異性。
- A. Correlation (相關係數): 用於衡量兩個變數之間的線性關係,而非單一變數的變異性。因此,最不恰當。
- B. Interquartile (四分位距): 衡量資料中間 50% 的離散程度,適用於比例尺度。
第 7 題
- Which of the following is the most appropriate statistic to summary the association between the data for the floor area (unit: square foot) and household income (unit: thousand dollars)?
A. Correlation
B. Interquartile
C. Mean
D. Relative frequency
E. Skewness
登入後即可作答並保存紀錄。
這題考驗對兩個數值型變數 (numerical variables) 之間關聯性評估的統計量選擇。
題目評估的變數是「樓地板面積」(比例尺度)與「家庭收入」(比例尺度)。
- A. Correlation (相關係數): 衡量兩個數值型變數之間線性關係的強度與方向,是最適合的統計量。
第 8 題
- Which of the following is the most appropriate statistic to examine the shape of the distribution for the floor area (unit: square foot)?
A. Correlation
B. Interquartile
C. Mean
D. Relative frequency
E. Skewness
登入後即可作答並保存紀錄。
這題考驗對描述分佈形狀 (shape of distribution) 的統計量選擇。
題目評估的是「樓地板面積」(比例尺度)的「分佈形狀」。
- A. Correlation (相關係數): 衡量兩個變數的關係,與單一變數分佈形狀無關。
- B. Interquartile (四分位距): 衡量離散程度,間接反映分佈寬度,非直接描述形狀。
第 9 題
- A survey is distributed to understand the commuting methods to NTPU. 270 students were responded and the data were summarized as follows:
| Level of students | Method of Commutes | ||
|---|---|---|---|
| Metro | Bus | Walk | |
| freshman | 70 | 20 | 65 |
| Sophomore | 50 | 40 | 25 |
Determine the probability of the students who commute to school using metro or bus.
A. 1/9
B. 2/9
C. 1/3
D. 4/9
E. 2/3
登入後即可作答並保存紀錄。
這題考驗對機率計算的理解,特別是「或」事件的機率。
總學生人數 N = 270。
搭乘捷運 (Metro) 的總人數 = 70 (freshman) + 50 (Sophomore) = 120。
搭乘公車 (Bus) 的總人數 = 20 (freshman) + 40 (Sophomore) = 60。
由於搭乘捷運和搭乘公車是互斥事件(一個學生主要通勤方式只能選一種),計算搭乘捷運或公車的機率,只需將兩者人數相加後除以總人數。
搭乘捷運或公車的總人數 = 120 + 60 = 180。
第 10 題
- (Continued Question 9.) Given the respondents who were freshman, determine the probability of the students who walked to school?
A. 1/6
B. 13/31
C. 1/3
D. 5/18
E. 13/18
登入後即可作答並保存紀錄。
這題考驗條件機率 (conditional probability) 的計算。
題目要求在「受訪者為 freshman」的條件下,計算「學生選擇步行 (Walk) 上學的機率」,即 。
在 freshman 中,總人數為:70 (Metro) + 20 (Bus) + 65 (Walk) = 155 人。
其中,選擇步行上學的 freshman 有 65 人。
第 11 題
- What of the following property is not appropriate to describe event A and event B are independent ?
A.
B.
C.
D.
E.
登入後即可作答並保存紀錄。
核心觀念
本題考「事件獨立」的定義及其等價條件。事件 與事件 獨立,定義為
若 、,亦可表示為
此外,由聯集公式:
代入獨立條件後得
題目中的 、 應按選項脈絡理解為條件機率 、;若原題確實省略條件機率符號,應視為排版簡寫。
解題方法
逐一比對各選項是否能描述獨立事件:
- 條件機率不因另一事件發生而改變,是獨立的等價條件。
- 交集機率等於機率乘積,是獨立的正式定義。
- 聯集公式代入獨立條件後,可得到選項 E。
- 表示兩事件互斥,通常與獨立不同。
因此,不適合描述獨立事件的是選項 C。
選項分析
A.
此式表示在事件 發生的條件下,事件 的機率沒有改變,正是 與 獨立的等價條件。
因此 A 正確。
B.
此式表示在事件 發生的條件下,事件 的機率沒有改變,也是獨立事件的等價條件。
因此 B 正確。
第 12 題
- A doctor is concerned about the relationship between blood pressure and sugar level. Among her patients, she classifies blood pressures as high, normal, or low and sugar level as normal and abnormal. She finds that (a) 15% have high blood pressure; (b) 20% have low blood pressure; (c) 18% have an abnormal sugar level; (d) of those with an abnormal sugar level, 35% have high blood pressure; and (e) of those with normal blood pressure, 12% have an abnormal sugar level. What percentage of her patients have a normal sugar level and low blood pressure?
A. 0.161
B. 0.176
C. 0.184
D. 0.920
E. 0.039
登入後即可作答並保存紀錄。
這題考驗對條件機率和聯合機率的計算,涉及多個事件的交集。我們可以使用表格法來整理資訊。
首先,定義事件:
- BP_H: High blood pressure (高血壓)
- BP_N: Normal blood pressure (正常血壓)
- BP_L: Low blood pressure (低血壓)
- SUGAR_A: Abnormal sugar level (異常血糖)
- SUGAR_N: Normal sugar level (正常血糖)
已知資訊:
(a)
(b)
(c)
(d)
(e)
我們需要計算 。
由 (a) 和 (b),我們可以推斷出 :
由於血壓只有高、正常、低三種情況,且是互斥的,總機率為 1。
由 (c),我們可以推斷出 :
血糖只有異常和正常兩種情況。
現在,利用條件機率公式 ,我們可以計算出一些聯合機率:
從 (d) :
從 (e) :
我們可以建立一個聯合機率表格:
| SUGAR_A (0.18) | SUGAR_N (0.82) | Total | |
|---|---|---|---|
| BP_H (0.15) | 0.063 | ? | 0.15 |
| BP_N (0.65) | 0.078 | ? | 0.65 |
| BP_L (0.20) | ? | ? | 0.20 |
| --------------- | ---------------- | ---------------- | ---------- |
| Total | 0.18 | 0.82 | 1.00 |