114 年 國立臺北大學都市計劃研究所《統計學》

📄 試題原卷 免費註冊後即可對照原始考卷 PDF免費註冊

第 1 題25 分

  1. (25%) In response to the construction of the Neihu MRT, the Taipei City Department of Transportation hopes to understand the commuting conditions and needs of Neihu residents through a telephone survey. The questionnaire covers basic demographic information, the type of commuting tools used, commuting costs (in NT dollars), and commuting time (in minutes). The basic demographic information includes age (in years), gender, and education level (1 = high school or below; 2 = university; 3 = graduate school or above). Exploratory data analysis (EDA) is an important statistical analysis procedure. Data visualization is one aspect of EDA.

The commonly used statistical figures include:
(A) Bar chart
(B) Box plot
(C) Histogram
(D) Line chart (折線圖)
(E) Pie chart
(F) QQ plot
(G) Scatter plot

Use the letters (A) - (G) to answer the following questions. Multiple figures may be possible for the following questions.

(1) Write down the appropriate figures to visualize the type of commuting tools used.
(2) Write down the appropriate figures to visualize commuting costs.
(3) Write down the appropriate figures to visualize the association between commuting costs and commuting time.
(4) Write down the appropriate figures to visualize the association between gender and commuting time.
(5) Write down the appropriate figures to visualize the association between gender and the type of commuting tools used.

🖼️ 本題含圖表,以下為原卷對應頁面:
原卷第 1 頁

登入後即可作答並保存紀錄。

這一題的完整詳解

本題主要在於考驗學生對於資料視覺化的基本觀念,以及不同類型變數適合使用的圖表。

核心概念:資料視覺化、變數類型與圖表選擇

(1) 視覺化「通勤工具的種類」:
通勤工具的種類是一個類別變數(例如:公車、捷運、自行車、汽車等)。對於類別變數,我們通常使用長條圖或圓餅圖來呈現其頻率分佈。

  • 長條圖 (Bar chart, A) 可以清楚比較不同類別的數量。
  • 圓餅圖 (Pie chart, E) 可以顯示各類別佔整體的比例。
    因此,適合的圖表是 (A) 和 (E)。

(2) 視覺化「通勤成本」:
通勤成本(以新台幣計價)是一個連續型數值變數。

  • 直方圖 (Histogram, C) 可以顯示數值資料的頻率分佈,了解其集中趨勢、離散程度和偏態。
  • 箱型圖 (Box plot, B) 可以顯示資料的四分位距、中位數、最大值、最小值,以及異常值,適合呈現單一數值變數的摘要統計。
    因此,適合的圖表是 (B) 和 (C)。

(3) 視覺化「通勤成本與通勤時間的關聯性」:
通勤成本和通勤時間都是連續型數值變數。要探討兩個連續變數之間的關聯性,最常用的圖形是散佈圖。

  • 散佈圖 (Scatter plot, G) 可以顯示兩變數之間的關係模式,例如線性、非線性、正相關、負相關或無關。
    因此,適合的圖表是 (G)。
🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 2 題25 分

  1. (25%) In response to the construction of the Neihu MRT, the Taipei City Department of Transportation hopes to understand the commuting conditions and needs of Neihu residents through a telephone survey. The questionnaire covers basic demographic information, the type of commuting tools used, commuting costs (in NT dollars), and commuting time (in minutes). The basic demographic information includes age (in years), gender, and education level (1 = high school or below; 2 = university; 3 = graduate school or above). Exploratory data analysis (EDA) is an important statistical analysis procedure. The followings are some commonly used summary statistics:

(A) Correlation
(B) Frequency
(C) Interquartile
(D) Kurtosis
(E) Mean
(F) Medium
(G) Minimum
(H) Mode
(I) Percent
(J) Skewness
(K) Standard deviation
(L) Variance

Use the letters (A) - (L) to answer the following questions. Multiple summary statistics might be possible for the following questions.

(1) Write down the appropriate summary statistics to summarize the type of commuting tools used.
(2) Write down the appropriate summary statistics to describe the central tendency of commuting costs.
(3) Write down the appropriate summary statistics to describe the variability of commuting costs.
(4) Write down the appropriate summary statistics to describe the shape of the distribution of commuting costs.
(5) Write down the appropriate summary statistics to summarize the association between commuting cost and commuting time.

🖼️ 本題含圖表,以下為原卷對應頁面:
原卷第 1 頁

登入後即可作答並保存紀錄。

這一題的完整詳解

本題主要在於考驗學生對於不同類型變數的描述性統計量以及衡量變數間關聯性的統計量之認識。

核心概念:描述性統計量、集中趨勢、離散趨勢、分佈形狀、變數間關聯性。

(1) 總結「通勤工具的種類」:
通勤工具的種類是一個類別變數。描述類別變數時,我們通常關注其頻率或比例。

  • 頻率 (Frequency, B):計算每個類別出現的次數。
  • 百分比 (Percent, I):計算每個類別佔總體的比例。
  • 眾數 (Mode, H):代表最常出現的類別。
    因此,適合的統計量是 (B), (I), (H)。

(2) 描述「通勤成本」的「集中趨勢」:
通勤成本是一個連續型數值變數。描述集中趨勢(平均值、代表值)的統計量有:

  • 平均數 (Mean, E):所有數值的總和除以數值個數。
  • 中位數 (Medium, F):將所有數值排序後,位於最中間的數值(若個數為偶數,則為中間兩個數值的平均)。
  • 眾數 (Mode, H):出現頻率最高的數值。
    因此,適合的統計量是 (E), (F), (H)。

(3) 描述「通勤成本」的「變異性」:
變異性(或稱為離散程度)描述數據點分散的程度。

🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 3 題15 分

  1. (15%) In response to the construction of the Neihu MRT, the Taipei City Department of Transportation hopes to understand the commuting conditions and needs of Neihu residents through a telephone survey. The questionnaire covers basic demographic information, the type of commuting tools used, commuting costs (in NT dollars), and commuting time (in minutes). The basic demographic information includes age (in years), gender, and education level (1 = high school or below; 2 = university; 3 = graduate school or above). Besides EDA, it is also important to identify possible association among variables. The following provides commonly used test statistics:

(A) Chi-square test
(B) Correlation
(C) McNemar test
(D) One-Way Analysis of Variance (ANOVA)
(E) Paired T test
(F) Simple linear regression
(G) Two independent sample T test
(H) Wilcoxon rank sum test

Use the letters (A) – (H) to answer the following questions. Multiple test statistics may be possible for the following questions.

(1) Write down the appropriate test statistics for testing the association between gender and the type of commuting tools used.
(2) Write down the appropriate test statistics for testing the association between commuting costs and age.
(3) Write down the appropriate test statistics for testing commuting time and gender.

🖼️ 本題含圖表,以下為原卷對應頁面:
原卷第 2 頁

登入後即可作答並保存紀錄。

這一題的完整詳解

本題主要在於考驗學生對於不同變數組合(類別與類別、數值與類別、數值與數值)適合使用的統計檢定方法之認識。

核心概念:統計檢定、獨立性檢定、相關性檢定、變異數分析、迴歸分析。

(1) 檢定「性別」與「通勤工具種類」的關聯性:
性別是一個類別變數,通勤工具種類也是一個類別變數。檢定兩個類別變數之間是否存在關聯性(或獨立性)的常用方法是:

  • 卡方獨立性檢定 (Chi-square test, A):用於檢定兩個類別變數之間是否獨立。
  • 雖然 Wilcoxon rank sum test (H) 適用於兩個獨立樣本的比較,但它通常用於比較兩個獨立樣本的中心趨勢(通常是連續變數),而不是兩個類別變數的關聯性。
    因此,最適合的檢定統計量是 (A)。

(2) 檢定「通勤成本」與「年齡」的關聯性:
通勤成本是一個連續型數值變數,年齡是一個連續型數值變數(雖然年齡常被視為離散,但在統計分析中常被當作連續變數處理,特別是當樣本數較大時)。

  • 相關係數 (Correlation, B):用於檢定兩個連續變數之間是否存在線性關聯。
  • 簡單線性迴歸 (Simple linear regression, F):可以檢定一個連續變數是否能預測另一個連續變數,並同時評估關聯性。
  • One-Way ANOVA (D) 通常用於比較三個或以上類別變數組別的連續變數均值。
🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 4 題10 分

  1. (10%) Let X denote commuting costs. Assume that the distribution of X approximately follows a normal distribution with mean 30 NT dollars and standard deviation 2 NT dollars.

(1) Find the probability that commuting costs exceeds 32 NT dollars.
(2) Assume the expert would like to know the average commuting costs for Neihu residents. The expert would like to collect a sample data to estimate the mean commuting costs μ. Assume the standard deviation equals 2 NT dollars. How large a sample is necessary if he want the estimate to be within 0.28 NT dollar of the actual mean value μ, with 95% confidence?

🖼️ 本題含圖表,以下為原卷對應頁面:
原卷第 2 頁

登入後即可作答並保存紀錄。

這一題的完整詳解

本題考驗學生對常態分佈機率計算以及樣本數估計的應用。

核心概念:常態分佈、標準化、機率計算、信賴區間、樣本數估計。

(1) 計算通勤成本超過 32 新台幣的機率:
已知通勤成本 X 服從常態分佈,其平均數 μ=30\mu = 30 NT dollars,標準差 σ=2\sigma = 2 NT dollars。
我們需要計算 P(X>32)P(X > 32)。
首先,將 X 標準化為 Z 分數:
Z=X−μσZ = \frac{X - \mu}{\sigma}
當 X=32X = 32 時,對應的 Z 分數為:
Z=32−302=22=1Z = \frac{32 - 30}{2} = \frac{2}{2} = 1
因此,P(X>32)=P(Z>1)P(X > 32) = P(Z > 1)。
查標準常態分佈表,P(Z ≤ 1) 約為 0.8413。
所以,P(Z>1)=1−P(Z≤1)=1−0.8413=0.1587P(Z > 1) = 1 - P(Z \le 1) = 1 - 0.8413 = 0.1587。

(2) 計算所需的樣本數:
專家希望估計平均通勤成本 μ\mu,並要求估計值與真實值 μ\mu 的誤差在 0.28 NT dollar 以內,且信賴水準為 95%。

🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 5 題15 分

  1. (15%) Assume the target commuting costs equals 30 NT dollars. Suppose a sample of size 25 residents are asked and assume the probability that the residents would be willing to pay more than 30 NT dollars equal 0.1. Let X equal the number of residents who pay more than 30 NT dollars to commute.

(1) Find P[X = 1].
(2) Find the probability that no residents pay more than 30 NT dollars for their daily commute.
(3) Find mean and variance of X.

🖼️ 本題含圖表,以下為原卷對應頁面:
原卷第 2 頁

登入後即可作答並保存紀錄。

這一題的完整詳解

本題考驗學生對二項分佈的理解與應用,包含機率計算、期望值與變異數的求解。

核心概念:二項分佈、機率質量函數、期望值、變異數。

題目說明:

  • 樣本大小 n=25n = 25。
  • 每位居民願意支付超過 30 NT dollars 的機率為 p=0.1p = 0.1。
  • X 是願意支付超過 30 NT dollars 的居民人數。
    這符合二項分佈的條件:固定次數的獨立重複試驗,每次試驗只有兩種結果(成功:願意支付超過 30 NT dollars,失敗:不願意支付超過 30 NT dollars),且成功的機率固定。
    因此,X 服從二項分佈,記為 X∼B(n,p)X \sim B(n, p),即 X∼B(25,0.1)X \sim B(25, 0.1)。
    二項分佈的機率質量函數為:P(X=k)=(nk)pk(1−p)n−kP(X = k) = \binom{n}{k} p^k (1-p)^{n-k}。

(1) 計算 P[X = 1]:
此處 k=1k=1。
P(X=1)=(251)(0.1)1(1−0.1)25−1P(X = 1) = \binom{25}{1} (0.1)^1 (1-0.1)^{25-1}
P(X=1)=25⋅(0.1)⋅(0.9)24P(X = 1) = 25 \cdot (0.1) \cdot (0.9)^{24}
P(X=1)=2.5⋅(0.9)24P(X = 1) = 2.5 \cdot (0.9)^{24}
計算 (0.9)24(0.9)^{24}:
(0.9)24≈0.079766(0.9)^{24} \approx 0.079766
P(X=1)≈2.5⋅0.079766≈0.199415P(X = 1) \approx 2.5 \cdot 0.079766 \approx 0.199415

🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

第 6 題10 分

  1. (10%) The Taipei City Department of Transportation expects that there exists a gender difference in commuting time. Assume that the distributions of the commuting time for female and male are normally distributed. Let the sample mean and sample deviation of commuting time for female and male be:
nSample meanSample varianceSample standard deviation
Female2553.35100.4110.02
Male2555.53112.2010.59

Assume that the variance of the commuting time are roughly equal between female and male.
(1) Write down the hypotheses. The proper notation should be used.
(2) Use a 5% significance level to make your inference.

🖼️ 本題含圖表,以下為原卷對應頁面:
原卷第 2 頁原卷第 3 頁

登入後即可作答並保存紀錄。

這一題的完整詳解

本題考驗學生對獨立樣本 t 檢定(未混合變異數)的應用,包含建立假設、計算檢定統計量、決定是否拒絕虛無假設。

核心概念:獨立樣本 t 檢定、虛無假設、對立假設、檢定統計量、顯著水準、p-value 或臨界值判斷。

題目假設:

  • 女性與男性的通勤時間分別服從常態分佈。
  • 兩組的變異數大致相等(此為進行聯合變異數估計的基礎)。
  • 樣本大小:nF=25n_F = 25, nM=25n_M = 25。
  • 樣本均值:xˉF=53.35\bar{x}_F = 53.35, xˉM=55.53\bar{x}_M = 55.53。
  • 樣本變異數:sF2=100.41s_F^2 = 100.41, sM2=112.20s_M^2 = 112.20。
  • 樣本標準差:sF=10.02s_F = 10.02, sM=10.59s_M = 10.59。
  • 顯著水準 α=0.05\alpha = 0.05。

(1) 建立假設:
我們要檢定的是性別之間通勤時間是否存在差異。

  • 虛無假設 (H0H_0):女性與男性的平均通勤時間沒有差異。
    μF=μM\mu_F = \mu_M 或 μF−μM=0\mu_F - \mu_M = 0
  • 對立假設 (H1H_1):女性與男性的平均通勤時間有差異(雙尾檢定)。
    μF≠μM\mu_F \neq \mu_M 或 μF−μM≠0\mu_F - \mu_M \neq 0

(2) 進行推論(使用 5% 顯著水準):
由於假設兩組變異數大致相等,我們採用合併變異數的獨立樣本 t 檢定。

步驟一:計算合併變異數 (sp2s_p^2)。
合併變異數的公式為:
sp2=(nF−1)sF2+(nM−1)sM2nF+nM−2s_p^2 = \frac{(n_F - 1)s_F^2 + (n_M - 1)s_M^2}{n_F + n_M - 2}
sp2=(25−1)⋅100.41+(25−1)⋅112.2025+25−2s_p^2 = \frac{(25 - 1) \cdot 100.41 + (25 - 1) \cdot 112.20}{25 + 25 - 2}
sp2=24⋅100.41+24⋅112.2048s_p^2 = \frac{24 \cdot 100.41 + 24 \cdot 112.20}{48}
sp2=2409.84+2692.848s_p^2 = \frac{2409.84 + 2692.8}{48}
sp2=4702.6448≈97.9717s_p^2 = \frac{4702.64}{48} \approx 97.9717

步驟二:計算 t 檢定統計量 (tt)。

🔒

後續完整解題步驟與【答案】

免費註冊,享三天全站完整詳解閱覽。

免費註冊

其他考古題