111 年 國立臺北大學都市計劃研究所甲組《統計學》
第 1 題25 分
- (25%) Consider a data for investigating the housing price (unit: thousand dollars) in an area. Five attributes were collected including the location (metropolitan area, urban and suburban), floor size (units: feet), distance to the metro station (units: kilometer), number of bathrooms (1, 2, 3) and having the balcony (yes and no). Exploratory data analysis (EDA) is an important statistical analysis procedure. Data visualization is one of the EDA. The following provides common used statistical figures:
(A) 🖼️【此處有附圖,請對照原卷】
(B) 🖼️【此處有附圖,請對照原卷】
(C) 🖼️【此處有附圖,請對照原卷】
(D) 🖼️【此處有附圖,請對照原卷】
(E) 🖼️【此處有附圖,請對照原卷】
(F) 🖼️【此處有附圖,請對照原卷】
(G) 🖼️【此處有附圖,請對照原卷】
(H) 🖼️【此處有附圖,請對照原卷】
(I) 🖼️【此處有附圖,請對照原卷】
Use the letters A - I to answer the following questions. Multiple figures might be possible for the following questions.
(i) Write down the appropriate figures for visualizing distance to the metro station.
(ii) Write down the appropriate figures for visualizing location.
(iii) Write down the appropriate figures for visualizing location and housing price.
(iv) Write down the appropriate figures for visualizing distance to the metro station and housing price.
(v) Write down the appropriate figures for visualizing having balcony and the location.
登入後即可作答並保存紀錄。
本題主要在考察資料視覺化(Data Visualization)的應用,以及不同圖表類型適合呈現的資料特性與關係。
核心觀念:
- 單變數視覺化:如何呈現單一變數的分布或特徵。
- 多變數視覺化:如何呈現兩個或多個變數之間的關係。
- 變數類型:定性變數(如 location, balcony)與定量變數(如 distance, floor size, housing price)的視覺化方法。
題目分析與解題思路:
我們需要根據每個子題所要視覺化的變數類型及其個數,來選擇最適合的圖表。
變數說明:
- Housing Price: 定量 (連續)
- Location: 定性 (三個類別:metropolitan area, urban, suburban)
- Floor Size: 定量 (連續)
- Distance to metro station: 定量 (連續)
- Number of bathrooms: 定性 (三個類別:1, 2, 3)
- Having balcony: 定性 (兩個類別:yes, no)
圖表類型判斷:
- 長條圖 (Bar Chart):適合呈現定性變數的頻率分布,或離散定量變數的分布。圖 (A) 似乎是長條圖。
- 盒鬚圖 (Box Plot):適合呈現定量變數的分布摘要(中位數、四分位數、離群值),常用於比較不同群組間的分布。圖 (B) 可能是盒鬚圖。
- 直方圖 (Histogram):適合呈現定量變數的頻率分布。圖 (C) 似乎是直方圖。
- 散點圖 (Scatter Plot):適合呈現兩個定量變數之間的關係。圖 (D) 似乎是散點圖。
- 折線圖 (Line Chart):通常用於顯示隨時間變化的趨勢,或兩個定量變數的關係,尤其當其中一個變數(如 X 軸)是連續且有序時。圖 (E), (F), (G) 可能是折線圖或類似圖表。
- 圓餅圖 (Pie Chart):適合呈現定性變數各類別的比例,但通常不如長條圖清晰。圖 (H) 可能是圓餅圖。
- 散點圖矩陣 (Scatter Plot Matrix):顯示多個變數對兩兩之間的散點圖。圖 (I) 可能是某種多變數圖。
詳細解答:
(i) Write down the appropriate figures for visualizing distance to the metro station.
「Distance to the metro station」是一個單一的定量變數。適合用來呈現單一定量變數分布的圖表有直方圖和盒鬚圖。
- 直方圖 (Histogram):顯示數值的分布情況,例如集中趨勢、離散程度、偏態等。
- 盒鬚圖 (Box Plot):顯示中位數、四分位數、極端值等統計摘要,適合觀察分布的集中趨勢、離散程度和潛在的離群值。
因此,圖 (C)(直方圖)和圖 (B)(盒鬚圖)都是適合的。
【答案】(C), (B)
(ii) Write down the appropriate figures for visualizing location.
「Location」是一個定性變數(metropolitan area, urban, suburban)。
第 2 題20 分
- (20%) Consider a data for investigating the housing price (unit: thousand dollars) in an area. Five attributes were collected including the location (metropolitan area, urban and suburban), floor size (units: feet), distance to the metro station (units: kilometer), number of bathrooms (1, 2, 3) and having the balcony (yes and no). The first step for the data analyses is to compute descriptive statistics for each variables. The following provides common used summary statistics:
(A) Correlation
(B) Frequency
(C) Interquartile
(D) Kurtosis
(E) Mean
(F) Medium
(G) Minimum
(H) Mode
(I) Percent
(J) Skewness
(K) Standard deviation
(L) Variance
Use the letters A - L to answer the following questions. Multiple summary statistics might be possible for the following questions.
(i) Write down the appropriate summary statistics for describing the centrality of the distance to the metro station.
(ii) Write down the appropriate summary statistics for describing the variability of the distance to the metro station.
(iii) Write down the appropriate summary statistics for describing the shape of the distribution of the distance to the metro station.
(iv) Write down the appropriate summary statistics for summarizing the location.
登入後即可作答並保存紀錄。
本題主要在考察敘述統計學(Descriptive Statistics)中,用來描述資料集中趨勢、離散程度及分布形狀的各種統計量。
核心觀念:
- 集中趨勢 (Centrality):描述資料的主要數值在哪裡。
- 離散趨勢 (Variability/Dispersion):描述資料的散布範圍或變異程度。
- 分布形狀 (Shape):描述資料的偏態(Skewness)和峰態(Kurtosis)。
- 變數類型:不同的統計量適用於不同類型的變數(定量 vs. 定性)。
題目分析與解題思路:
我們需要根據子題所問的「描述目的」以及「變數類型」,來選擇最適合的敘述統計量。
變數說明:
- Distance to the metro station: 定量 (連續)。
- Location: 定性 (三個類別:metropolitan area, urban, suburban)。
統計量列表:
- 集中趨勢: Mean (E), Median (F), Mode (H), Minimum (G), Percent (I - in some contexts, like percentiles)
- 離散趨勢: Standard deviation (K), Variance (L), Interquartile range (C)
- 分布形狀: Kurtosis (D), Skewness (J)
- 其他: Frequency (B) - for counts, Correlation (A) - for relationship between two quantitative variables.
詳細解答:
(i) Write down the appropriate summary statistics for describing the centrality of the distance to the metro station.
「Distance to the metro station」是一個定量變數。描述集中趨勢的統計量主要有:
- Mean (E):平均值,是最常見的集中趨勢指標。
- Median (F):中位數,對離群值較不敏感。
- Mode (H):眾數,指出現次數最多的數值,在連續數據中較少單獨使用,但若數據分組則可出現。
- Minimum (G):最小值,雖然是極端值,但也可視為集中趨勢的一個極端點。
因此,Mean, Median, Mode 都是描述集中趨勢的統計量。
【答案】(E), (F), (H)
第 3 題15 分
- (15%) Consider a data for investigating the housing price (unit: thousand dollars) in an area. Five attributes were collected including the location (metropolitan area, urban and suburban), floor size (units: feet), distance to the metro station (units: kilometer), number of bathrooms (1, 2, 3) and having the balcony (yes and no). The second step for the data analyses is to find the bivariate association between housing price and other attributes. The following provides common used test statistics:
(A) One-Way Analysis of Variance (ANOVA)
(B) Correlation
(C) Chi-square test
(D) McNemar test
(E) Paired T test
(F) Simple linear regression
(G) Two independent sample T test
(H) Wilcoxon rank sum test
Use the letters A - H to answer the following questions. Multiple test statistics might be possible for the following questions.
(i) Write down the appropriate test statistics for testing the association between the housing price and location.
(ii) Write down the appropriate test statistics for testing the association between the housing price and the floor size.
(iii) Write down the appropriate test statistics for testing the association between the housing price and having the balcony.
登入後即可作答並保存紀錄。
本題主要在考察如何選擇適當的統計檢定方法,來分析不同變數類型之間的關聯性(bivariate association)。
核心觀念:
- 變數類型與檢定方法:
- 定量 vs. 定量:相關分析 (Correlation)、迴歸分析 (Regression)。
- 定量 vs. 定性(兩組):獨立樣本 t 檢定 (Independent samples t-test)。
- 定量 vs. 定性(多組):單因子變異數分析 (One-way ANOVA)。
- 定性 vs. 定性:卡方檢定 (Chi-square test)。
- 檢定目的:是檢定「關聯性」還是「差異性」。
題目分析與解題思路:
我們需要根據子題所要檢定的兩個變數的類型,來選擇最適合的檢定統計量。
變數說明:
- Housing Price: 定量 (連續)
- Location: 定性 (三個類別:metropolitan area, urban, suburban)
- Floor Size: 定量 (連續)
- Having balcony: 定性 (兩個類別:yes, no)
常見檢定方法對應的變數類型:
- (A) One-Way ANOVA:用於檢定一個定性變數(多於兩組)與一個定量變數之間是否存在顯著差異。
- (B) Correlation:用於檢定兩個定量變數之間是否存在線性關聯。
- (C) Chi-square test:用於檢定兩個定性變數之間是否存在關聯。
- (D) McNemar test:用於檢定配對樣本的兩個定性變數之間是否存在關聯(常用於前後測或配對設計)。
- (E) Paired T test:用於檢定配對樣本的兩個定量變數之間是否存在顯著差異。
- (F) Simple linear regression:用於建立一個定量變數(應變數)與一個定量變數(自變數)之間的線性模型,並檢定自變數的顯著性。
- (G) Two independent sample T test:用於檢定兩個獨立樣本(或一個定性變數兩組)與一個定量變數之間是否存在顯著差異。
- (H) Wilcoxon rank sum test:獨立樣本 t 檢定的非參數版本,用於檢定兩個獨立樣本(或一個定性變數兩組)與一個定量變數之間是否存在顯著差異。
詳細解答:
(i) Write down the appropriate test statistics for testing the association between the housing price and location.
- Housing Price:定量
- Location:定性(三個類別:metropolitan area, urban, suburban)
我們要檢定的是一個「多組定性變數」與「一個定量變數」之間的關聯性。
- One-Way ANOVA (A):正是用於檢定一個分組變數(多於兩組)與一個定量反應變數之間是否存在顯著差異。
- Correlation (B):不適用,因為 Location 是定性變數。
- Chi-square test (C):不適用,因為 Housing Price 是定量變數。
第 4 題10 分
- (10%) Let denote the housing price. Assume that the distribution of approximately follows a normal distribution with mean 75 millions and standard deviation 100 millions.
(i) Find the probability that the housing price exceeds 90 millions.
(ii) Assume a businessman would like to know the reasonable housing price. He would like to conduct a study to collect a sample to estimate the mean housing price . Assume the standard deviation equals 100 millions. How large a sample is necessary if he want the estimate to be within 2 millions of the actual mean value , with 95% confidence?
登入後即可作答並保存紀錄。
本題主要在考察常態分布的機率計算以及樣本大小的決定。
核心觀念:
- 常態分布機率計算:利用標準化(Z-score)將常態分布轉換為標準常態分布,再利用標準常態分布表(或軟體)查詢機率。
- 樣本大小決定:根據預期的誤差範圍 (margin of error) 和信心水準 (confidence level),計算出所需的最小樣本大小。
題目分析與解題思路:
(i) Find the probability that the housing price exceeds 90 millions.
已知:
- 百萬
- 百萬
- 求
解題步驟:
- 標準化:將 轉換為標準常態變數 。
- 計算機率:求 。
由於標準常態分布表通常給出 ,我們利用機率的互補性:
- 查表:查標準常態分布表,找到 對應的累積機率 。
從提供的標準常態分布表(Page 4)中,當 時, (查 0.1 行, 0.05 列)。
計算:
【答案】
第 5 題10 分
- (10%) Assume a regular condo price at a given area is about 100 thousand dollars. Suppose the probability that the condo price exceeds 150 thousand dollars is 0.1. Let equal the number of condo prices that exceeds 150 thousand dollars in a sample of 30 condos selected at random from this area.
(i) Find .
(ii) Find mean and variance of .
登入後即可作答並保存紀錄。
本題主要考察二項分布(Binomial Distribution)的機率計算、期望值與變異數。
核心觀念:
- 二項分布的條件:
- 試驗次數固定 ()。
- 每次試驗只有兩種結果(成功/失敗)。
- 每次試驗的成功機率 () 相同。
- 各次試驗之間相互獨立。
- 二項分布機率公式:
- 二項分布的期望值:
- 二項分布的變異數:
題目分析與解題思路:
題目描述了一個抽樣情境,我們需要判斷其是否符合二項分布,並進行相關計算。
變數定義:
- 我們關注的「事件」(成功)是「 condo price exceeds 150 thousand dollars」。
- 已知「成功」的機率 。
- 樣本大小 (抽樣 30 個 condo)。
- 是在 30 個 condo 中,價格超過 15 萬美元的 condo 的數量。
此情境符合二項分布的條件:
- 試驗次數固定:。
- 結果兩種:超過 15 萬美元 (成功) 或不超過 (失敗)。
- 成功機率固定:。
- 假設抽樣是獨立的。
因此,。
詳細解答:
第 6 題10 分
- (10%) The greatest variation in real estate price is in California, where recently there was a difference of more than 1.2 million in 2001, whereas a similar home in Bakersfield was appraised at \mu1.2 million in 2002?
(i) Write down the hypotheses.
(ii) Use a 5% significance level to make your inference.
登入後即可作答並保存紀錄。
本題主要在考察假設檢定(Hypothesis Testing)的應用,特別是關於平均數的檢定,以及如何根據信賴區間來判斷檢定結果。
核心觀念:
- 假設檢定步驟:
- 建立虛無假設 () 與對立假設 ()。
- 選擇統計檢定方法與顯著水準 ()。
- 計算檢定統計量。
- 決定檢定區域(拒絕域)。
- 根據檢定統計量做出決策(拒絕或不拒絕 )。
- 信賴區間與假設檢定關係:
- 如果一個值(例如,虛無假設中的母體參數值)落在給定的信賴區間內,則在對應的顯著水準下,我們不拒絕虛無假設。
- 反之,如果該值落在信賴區間外,則我們拒絕虛無假設。
題目分析與解題思路:
題目給了一個信賴區間,並要求進行假設檢定。這是典型的「信賴區間法」應用。
已知資訊:
- 抽樣對象:Palo Alto 的房屋。
- 樣本大小:。
- 假設:房屋價格服從常態分布。
- 95% 信賴區間: 百萬美元。
- 欲檢定的值:平均價格 是否等於 1.2 百萬美元。
- 顯著水準:。
詳細解答:
(i) Write down the hypotheses.
我們要檢定的是「Palo Alto 的平均房屋價格 是否不同於 1.2 百萬美元」。
- 虛無假設 ():代表沒有改變或沒有差異,即平均價格等於 1.2 百萬美元。
第 7 題10 分
- (10%) Winthrop Boat Lines is exploring the possibility of offering a ferry service between the cities of Patna and Madura, provided there is sufficient demand to make it feasible. The firm randomly interviewed 210 commuters from the two cities, and 146 of them indicated they would patronize the ferry service instead of the present bus service.
(i) Estimate the population p of commuters from the two cities who would prefer the ferry service.
(ii) Construct a 95% confidence interval for p.
登入後即可作答並保存紀錄。
本題主要考察比例(Proportion)的點估計與區間估計。
核心觀念:
- 比例的點估計:樣本比例 是母體比例 的最佳點估計。
- 比例的信賴區間:當樣本量足夠大時,母體比例的信賴區間可以用近似常態分布的方法計算。
信賴區間公式:
其中 是對應於所選信心水準的標準常態值。
題目分析與解題思路:
題目給了一個抽樣調查的結果,要求估計總體的比例並建構信賴區間。
已知資訊:
- 總樣本數:。
- 表示願意搭乘渡輪的樣本數:。
- 信心水準:95%。
- 顯著水準:。
詳細解答:
(i) Estimate the population p of commuters from the two cities who would prefer the ferry service.
母體比例 的點估計是樣本比例 。
代入數值:
計算 :
因此,估計的母體比例約為 0.6952。
【答案】
(ii) Construct a 95% confidence interval for p.
我們需要建立一個 95% 的信賴區間。首先,我們需要計算對應於 95% 信心水準的 值。
- (從標準常態分布表中查得)