113 年 國立臺北大學都市計劃研究所甲組《統計學》
第 1 題25 分
An urban planning expert would like to understand the need of the public transportation in a residential area. A questionnaire is designed to collect the residents' opinions. The question in the questionnaire includes basic demographic variables, the type of public transportation and the willingness to pay fare (unit: NT dollar). The basic demographic variables include gender, age (unit: year) and type of jobs and years of education. Exploratory data analysis (EDA) is an important statistical analysis procedure. Data visualization is one of the EDA.
The commonly used statistical figures include:
(A) Bar chart
(B) Box plot
(C) Histogram
(D) Line chart
(E) Pie chart
(F) QQ plot
(G) Scatter plot
Use the letters (A) – (G) to answer the following questions. Multiple figures might be possible for the following questions.
(i) Write down the appropriate figures for visualizing the willingness to pay fare.
(ii) Write down the appropriate figures for visualizing gender.
(iii) Write down the appropriate figures for visualizing the type of jobs.
(iv) Write down the appropriate figures for visualizing sex and willingness to pay fare.
(v) Write down the appropriate figures for visualizing age and willing to pay.
登入後即可作答並保存紀錄。
本題主要在考資料視覺化(Data Visualization)的應用,目的是藉由圖形來呈現資料的分布與關係。
核心觀念:
不同的資料類型(連續、類別)和分析目的(單變數分布、雙變數關係)適合不同的統計圖形。
解題思路:
-
理解資料類型:
- 意願支付車費(willingness to pay fare):連續型變數(NT dollar)。
- 性別(gender):類別型變數(例如:男、女)。
- 工作類型(type of jobs):類別型變數(例如:學生、上班族、退休等)。
- 年齡(age):連續型變數(year),但常被分組為類別型(例如:青年、中年、老年)來分析。
-
理解圖形功能:
- (A) Bar chart (長條圖):適合呈現類別資料的頻率或比例,或比較離散數值。
- (B) Box plot (盒鬚圖):適合呈現連續資料的分布概況(中位數、四分位數、極值、異常值),特別適合比較不同群組的連續資料。
- (C) Histogram (直方圖):適合呈現連續資料的分布形狀(頻率分布)。
- (D) Line chart (折線圖):適合呈現時間序列資料的趨勢,或連續變數隨另一個連續變數變化的趨勢(但在本題的情境下較少用,除非是年齡與車費的趨勢圖)。
- (E) Pie chart (圓餅圖):適合呈現類別資料各部分佔整體的比例,但通常只適用於少數類別。
- (F) QQ plot (Q-Q 圖):適合檢定資料是否符合某種理論分布(如常態分布)。
- (G) Scatter plot (散佈圖):適合呈現兩個連續變數之間的關係(關聯性)。
-
逐一回答小題:
(i) 意願支付車費(willingness to pay fare):
這是一個連續變數,我們想了解其分布。
* 直方圖 (C) 最適合用來視覺化單一連續變數的分布形狀。
* 盒鬚圖 (B) 適合用來呈現單一連續變數的摘要統計(中位數、四分位數、離散程度),也可作為比較之用。
* 因此,(C) 和 (B) 都是合適的。
第 2 題25 分
An urban planning expert would like to understand the need of the public transportation in a residential area. A questionnaire is designed to collect the residents' opinions. The question in the questionnaire includes basic demographic variables, the type of public transportation and the willing to pay fare (unit: NT dollar). The basic demographic variables include gender, age (unit: year) and type of jobs and years of education. Exploratory data analysis (EDA) is an important statistical analysis procedure. The following provides the commonly used summary statistics.
(A) Correlation
(B) Frequency
(C) Interquartile
(D) Kurtosis
(E) Mean
(F) Median
(G) Minimum
(H) Mode
(I) Percent
(J) Skewness
(K) Standard deviation
(L) Variance
Use the letters (A) - (L) to answer the following questions. Multiple summary statistics might be possible for the following questions.
(i) Write down the appropriate summary statistics for the type of jobs.
(ii) Write down the appropriate summary statistics for describing the central tendency of the willingness to pay fare.
(iii) Write down the appropriate summary statistics for describing the variability of the willingness to pay fare.
(iv) Write down the appropriate summary statistics for describing the shape of the distribution of the willingness to pay fare.
(v) Write down the appropriate summary statistics for summarizing the association between the age and the willingness to pay fare.
登入後即可作答並保存紀錄。
本題主要在考各種統計量(summary statistics)的適用性,以及如何用它們來描述資料的特徵。
核心觀念:
不同的統計量有不同的用途,用於描述資料的集中趨勢、離散程度、分布形狀或變數間的關係。
解題思路:
-
理解統計量功能:
- (A) Correlation (相關係數):衡量兩個連續變數之間線性關聯的強度與方向。
- (B) Frequency (頻率):類別資料或離散資料中,特定類別或數值出現的次數。
- (C) Interquartile (四分位數):將資料分成四等分的點,通常指 IQR (Interquartile Range)。IQR = Q3 - Q1,衡量中間 50% 資料的離散程度。
- (D) Kurtosis (峰度):衡量分布的尖峭或平坦程度,相較於常態分布。
- (E) Mean (平均數):所有數值的總和除以數值的個數,衡量集中趨勢。
- (F) Median (中位數):排序後位於中間的數值,衡量集中趨勢,對離群值不敏感。
- (G) Minimum (最小值):資料中的最小數值。
- (H) Mode (眾數):出現頻率最高的數值,適用於類別或連續資料。
- (I) Percent (百分比):頻率除以總數,表示比例。
- (J) Skewness (偏度):衡量分布的不對稱性。
- (K) Standard deviation (標準差):衡量資料離散程度,是變異數的平方根。
- (L) Variance (變異數):衡量資料離散程度,是各數值與平均數差的平方的平均值。
-
理解資料類型與分析目的:
- 工作類型(type of jobs):類別變數。
- 意願支付車費(willingness to pay fare):連續變數。
- 年齡(age):連續變數。
-
逐一回答小題:
(i) 工作類型(type of jobs):
這是一個類別變數。我們需要描述其組成。
* Frequency (B) 和 Percent (I) 是最適合用來描述類別變數的組成(例如:有多少人是學生、多少人是上班族,各佔多少比例)。
* Mode (H) 也可以用來找出最常見的工作類型。
* 因此,(B), (I), (H) 是合適的。(ii) 意願支付車費(willingness to pay fare)的集中趨勢(central tendency):
意願支付車費是連續變數。集中趨勢的衡量指標有:
* Mean (E):平均數。
* Median (F):中位數。
* Mode (H):眾數。
第 3 題15 分
An urban planning expert would like to understand the need of the public transportation in a residential area. A questionnaire is designed to collect the residents' opinions. The question in the questionnaire includes basic demographic variables, the type of public transportation and the willing to pay fare (unit: NT dollar). The basic demographic variables include gender, age (unit: year) and type of jobs and years of education. Exploratory data analysis (EDA) is an important statistical analysis procedure. The following provides common used test statistics:
(A) Chi-square test
(B) Correlation
(C) McNemar test
(D) One-Way Analysis of Variance (ANOVA)
(E) Paired T test
(F) Simple linear regression
(G) Two independent sample T test
(H) Wilcoxon rank sum test
Use the letters (A) – (H) to answer the following questions. Multiple test statistics might be possible for the following questions.
(i) Write down the appropriate test statistics for testing the association between the willingness to pay fare and gender.
(ii) Write down the appropriate test statistics for testing the association between the willingness to pay fare and age.
(iii) Write down the appropriate test statistics for testing the type of public transportation and gender.
登入後即可作答並保存紀錄。
本題旨在考驗考生對於不同統計檢定方法適用時機的理解,特別是針對不同類型變數之間的關聯性檢定。
核心觀念:
不同的統計檢定方法適用於檢定不同類型變數之間的關聯性。
解題思路:
-
理解變數類型:
- 意願支付車費 (willingness to pay fare):連續變數。
- 性別 (gender):類別變數(通常是二分類)。
- 年齡 (age):連續變數。
- 公共交通工具類型 (type of public transportation):類別變數。
-
理解統計檢定方法的功能:
- (A) Chi-square test (卡方檢定):主要用於檢定兩個類別變數之間是否獨立(無關聯)。
- (B) Correlation (相關係數):主要用於衡量兩個連續變數之間的線性關聯強度。
- (C) McNemar test (麥克尼馬檢定):用於檢定配對樣本的兩個類別變數之間是否存在差異,常用於配對設計的二分類變數。
- (D) One-Way ANOVA (單因子變異數分析):用於檢定一個類別變數(因子,至少三類)和一個連續變數之間是否存在顯著差異。
- (E) Paired T test (配對樣本 t 檢定):用於檢定配對樣本的兩個連續變數之間是否存在差異。
- (F) Simple linear regression (簡單線性迴歸):用於建立一個連續自變數與一個連續應變數之間的線性關係模型,並可進行預測。檢定自變數對應變數的影響。
- (G) Two independent sample T test (獨立樣本 t 檢定):用於檢定兩個獨立樣本的連續變數之間是否存在差異。
- (H) Wilcoxon rank sum test (威爾康森等級和檢定):無母數檢定,用於檢定兩個獨立樣本的連續變數(或排序變數)之間是否存在差異,相當於獨立樣本 t 檢定的無母數版本。
-
逐一回答小題:
第 4 題10 分
Let X denote the willingness to pay fare. Assume that the distribution of X approximately follows a normal distribution with mean 25 NT dollars and standard deviation 2 NT dollars.
(i) Find the probability that the willingness to pay fare exceeds 28 NT dollars.
(ii) Assume the expert would like to know the reasonable willingness to pay fare. The expert would like to collect a sample data to estimate the mean housing price μ. Assume the standard deviation equals 2 NT dollars. How large a sample is necessary if he want the estimate to be within 0.25 NT dollar of the actual mean value μ, with 95% confidence?
登入後即可作答並保存紀錄。
本題包含兩部分:第一部分為常態分布的機率計算,第二部分為樣本大小的決定。
核心觀念:
- 標準常態分布的機率計算。
- 樣本大小(sample size)的決定,以達到預設的信賴區間寬度。
解題思路:
(i) Probability Calculation
- 已知資訊:
- X 服從常態分布:
- 平均數 NT dollars
- 標準差 NT dollars
- 問題:求 。
- 步驟:
- 將 X 標準化,轉換為標準常態變數 Z。標準化公式為 。
- 計算對應於 的 Z 值:
- 將問題轉換為標準常態分布的機率:。
- 利用標準常態分布表(通常提供 的值)來計算。
。 - 查詢標準常態分布表,找到 的值。從 page 3 的 Table 1,當 z = 1.50 時,。
- 計算最終機率:
。
(ii) Sample Size Determination
- 已知資訊:
- 估計的母體平均數 。
- 母體標準差 NT dollars (注意:這裡假設母體標準差已知,通常在樣本大小決定時會使用母
第 5 題15 分
Assume the target fare equals 25 NT dollars. Suppose a sample of size 25 residents are asked and assume the probability that the residents would be willing to pay more 27 NT dollars equal 0.1. Let X equal the number of residents who are willing to pay more than 27 NT dollars.
(i) Find .
(ii) Find the probability that no residents are willing to pay more than 27 NT dollars.
(iii) Find mean and variance of X.
登入後即可作答並保存紀錄。
本題主要考驗對二項分布(Binomial Distribution)的理解與應用。
核心觀念:
當實驗滿足以下條件時,可視為二項分布:
- 實驗由固定數目的獨立重複試驗組成(試驗次數 固定)。
- 每次試驗只有兩種可能結果:「成功」(Success)或「失敗」(Failure)。
- 每次試驗成功的機率 相同,失敗的機率為 。
- 隨機變數 是成功次數。
二項分布的機率質量函數 (PMF) 為:
,其中 。
期望值 (Mean):
變異數 (Variance):
解題思路:
-
確認變數 X 的分布:
- 我們抽取了一個樣本,樣本大小 。
- 我們關心的是「願意支付超過 27 NT dollars」這個事件。這個事件可以視為一次試驗的「成功」。
- 題目假設「願意支付超過 27 NT dollars」的機率為 。
- 每次抽樣是獨立的。
- 因此,隨機變數 (願意支付超過 27 NT dollars 的人數)服從二項分布,記為 ,其中 ,。
-
逐一回答小題:
(i) 求 :
這是求在 25 次試驗中,恰好有 1 次成功的機率。
使用二項分布的 PMF:
代入 , , , :
第 6 題10 分
The urban planning expert expects that there exists a gender difference in willingness to pay fare. Assume that the fare between female and male are normally distributed. Let the sample mean and sample standard deviation of the willingness to pay fare for female and male be:
| n | Sample mean | Sample standard deviation | |
|---|---|---|---|
| Female | 25 | 26.85 | 1.11 |
| Male | 25 | 25.00 | 1.02 |
Assume that the variance of the fare are roughly equal between female and male.
(i) Write down the hypotheses. The proper notation should be used.
(ii) Use a 5% significance level to make your inference.
登入後即可作答並保存紀錄。
核心觀念
本題比較兩個獨立母體平均數:
- 女性願付票價母體平均數:
- 男性願付票價母體平均數:
題目說明兩組資料近似常態分布,且兩母體變異數大致相等,因此採用「等變異數的兩獨立樣本 檢定」,又稱 pooled two-sample test。
檢定統計量為
其中 pooled sample variance 為
自由度為
(i) 假設設定
題目敘述專家預期女性與男性在願付票價上存在性別差異,並未指定哪一性別較高,因此採雙尾檢定。
令:
- :女性願付票價的母體平均數
- :男性願付票價的母體平均數
假設為
亦可寫成
(ii) 以 5% 顯著水準進行推論
第一步:整理樣本資訊
女性:
男性:
樣本平均數差為
第二步:計算合併樣本變異數
先計算兩組樣本變異數:
因此
所以
第三步:計算檢定統計量
在虛無假設下,假定母體平均數差為 ,故