112 年 國立臺灣大學財務金融研究所乙組《統計學(I)》
請用 2B 鉛筆作答於答案卡,並先詳閱答案卡上之「畫記說明」。
Part I. 統計與計量(50%,選擇題,答案可能不只一個選項,每題 5 分)
第 1 題5 分
- Given the function:
for and for some , we define
and
Which of the following is right?
(a) and .
(b) and .
(c) and .
(d) and .
(e) None of the above choices (a)-(d).
登入後即可作答並保存紀錄。
核心觀念
本題評量微積分在機率統計中的基礎工具運用,主要涵蓋兩個核心概念:
- Gamma 函數(Gamma Function)與指數分配(Exponential Distribution)的動差積分:
對於常數 與整數 ,利用 Gamma 函數公式: 亦可連結至母數為 之指數分配隨機變數 的期望值計算:、。 - 瑕積分(Improper Integral)與動差生成函數(MGF)型積分的收斂性:
當 時,指數項之乘積依然滿足收斂條件:
解題方法
1. 計算
給定 ,將 拆解為兩個單獨積分:
利用標準積分公式 :
- 當 :
- 當 :
將兩者相加,可得:
2. 計算
將被積函數合併為單一指數:
由於題目設定 ,因此衰減指數之係數 ,該瑕積分收斂:
選項分析
第 2 題5 分
- Let be a sequence of independently and -distributed random variables. Suppose that we utilize Chebyshev's inequality to establish the following result:
where
Which of the following is right?
(a) .
(b) .
(c) when and .
(d) , as .
(e) None of the above choices (a)-(d).
登入後即可作答並保存紀錄。
核心觀念
本題的核心在於利用**柴比雪夫不等式(Chebyshev's inequality)**推導樣本變異數估計量(已知平均數 )偏離真實母體變異數 的機率上界,考驗以下重要定理與性質:
-
常態分配與卡方分配的關係:
若 ,則標準化變數 。
其平方項滿足 。 -
自由度為 1 之卡方分配動差:
若 ,其期望值為 ,變異數為 。 -
柴比雪夫不等式(Chebyshev's Inequality):
設隨機變數 具有有限期望值 與變異數 ,則對任意常數 ,恆滿足:
解題方法
令隨機變數 。
步驟一:計算隨機變數 的期望值與變異數
已知 ,因此可改寫為:
其中 。由此可得各單項的動差:
進而計算 的期望值與變異數:
由於 彼此獨立,故各 亦互相獨立:
第 3 題5 分
- Suppose that is a continuous random variable with and , and is a discrete random variable with possible outcomes: . Denote , and , for . Which of the following is right?
(a) and .
(b) and .
(c) and .
(d) and .
(e) None of the above choices (a)-(d).
登入後即可作答並保存紀錄。
核心觀念
本題考查機率論與數理統計中的兩大核心定理:
- 全期望值定律(Law of Iterated Expectations, LIE / Law of Total Expectation):
- 全變異數定律(Law of Total Variance / Eve's Law):
或是利用變異數的基本定義式進行條件拆解:。
解題方法與推導
1. 期望值 的推導
根據全期望值定律,對離散隨機變數 進行條件期望值運算:
因為 的可能取值為 ,機率為 ,且條件期望值為 ,因此展開為離散隨機變數的期望值加權和:
2. 變異數 的推導
方法一:利用第二級動差展開
隨機變數 的變異數定義為:
利用全期望值定律求 :
由條件變異數公式可知:
整理得條件第二級動差:
代入 的運算式:
將此結果代回 ,即可得到:
方法二:利用全變異數定律驗證
其中:
- 組內變異期望值(Expected Within-group Variance):
第 4 題5 分
- Assume that , where and are two uncorrelated standard normal random variables. Which of the following is right?
(a) and .
(b) and .
(c) and .
(d) and .
(e) None of the above choices (a)-(d).
登入後即可作答並保存紀錄。
核心觀念
本題評量標準常態分配的高階動差(Higher Moments)與隨機變數期望值的代數性質:
-
標準常態隨機變數的各階動差:
設 ,由於常態分配以 0 為中心對稱,所有奇數階動差皆為 0,偶數階動差可由動差母函數或雙階乘公式求得: -
零相關與獨立性(Uncorrelated vs. Independent):
在計量經濟學與統計學考題中,若提及 為常態且不相關(uncorrelated),通常隱含其具備聯合常態分配(Bivariate Normal Distribution)。對於二元常態隨機變數而言,「不相關()」與「相互獨立(Independent)」互為充分必要條件。因此,對任意函數 與 ,均有 。
解題方法
令 且 ,則 。
步驟一:求
展開 :
分別計算各項之期望值:
-
計算 :
取期望值:
-
計算 :
因 與 同分配,同理可得: -
計算 :
由獨立性可知:
第 5 題5 分
- Suppose that is a sequence of independently and -distributed random variables. Let be the probability density function of . Also, let and be the maximum likelihood estimators for and , respectively, that are solved by maximizing the log-likelihood function:
Which of the following is right?
(a) and .
(b) .
(c) .
(d) if .
(e) None of the above choices (a)-(d).
登入後即可作答並保存紀錄。
核心觀念
本題聚焦於常態分配母數之最大概似估計(Maximum Likelihood Estimation, MLE)與費雪訊息矩陣(Fisher Information Matrix)之核心性質:
- 單一觀測值之常態對數概似函數(Log-likelihood function)與分數函數(Score function):
常態分配 的機率密度函數為: 其對數形式為: - 分數函數的一階動差性質(First Bartlett Identity):
在滿足正則條件(Regularity conditions)下,分數函數對母數取偏微分後的期望值恆等於 : - 費雪訊息(Fisher Information)與參數的正交性(Orthogonality):
母數向量 的費雪訊息矩陣元素為: 在常態分配中,位置母數 與尺度母數 是資訊正交的(Informationally orthogonal),其分數函數之交叉乘積期望值必為 。
解題方法
首先對單一觀測值的對數概似函數分別對 與 求一階偏微分:
- 對 的偏微分:
- 對 的偏微分(此處將變異數視為單一母數 微分):
接著利用標準化隨機變數 的性質:
- (對稱分配無偏態)
- (常態峰態),因此
即可快速計算各動差與變異數。
選項分析
-
(a) 選項正確:
樣本平均對數概似函數為:令一階微分條件為零(First-Order Conditions, FOC):
經二階微分矩陣檢驗(負定),求得之 確實為全域最大值點。因此 (a) 正確。
-
(b) 選項錯誤:
根據正則條件下分數函數的期望值性質:
第 6 題5 分
- Let be a sequence of independently and identically distributed random variables with and a finite variance . Assume that . Denote and . Which of the following is right?
(a) has the limiting distribution , as , in general.
(b) has the limiting distribution , as , when for some fixed .
(c) has the limiting distribution , as , when for some fixed .
(d) has the limiting distribution , as , when for some fixed .
(e) None of the above choices (a)-(d).
登入後即可作答並保存紀錄。
核心觀念
本題評量大樣本漸近理論(Asymptotic Theory),特別是:
- 非典型樣本統計量的漸近性質:本題的樣本平均數與變異數定義分母為 (非一般的 或 ),需先釐清其與標準樣本平均數 及樣本變異數 的漸近等價關係。
- 中央極限定理(Central Limit Theorem, CLT):
- 連續映射定理(Continuous Mapping Theorem, CMT)與 Slutsky's Theorem。
- 局部對立假說(Local Alternative / Pitman Drift):當母體平均數 隨樣本數變化,設定為 時,檢定統計量的漸近非常態分配(Non-central limiting distribution)。
解題方法
1. 標準統計量與題目定義統計量之關係
令標準的樣本平均數為 ,題目定義為:
由於 ,兩者在漸近行為上完全等價。
同樣地,考慮樣本變異數項 :
將其展開:
兩邊同除以 :
根據大數法則(WLLN):
因此:
2. 分子 的漸近分配推導
將 拆解為標準 CLT 形式:
由 CLT 知 ,且 :
- 情況一(一般情況, 為不隨 變動的常數):
- 若 :。
- 若 :,此時 會發散(Diverge),不具退化的極限常態分配。
- 情況二(局部對立假說,):
代入得:
第 7 題5 分
- Let be an matrix of Gaussian random variables with . Assume that is positive definite. Define the matrix:
Which of the following is right?
(a) is positive definite.
(b) .
(c) .
(d) .
(e) None of the above choices (a)-(d).
登入後即可作答並保存紀錄。
核心觀念
本題考查計量經濟學與多元迴歸分析中**投影矩陣(Projection Matrix / Hat Matrix)**的代數性質:
- 對稱性(Symmetry):
若 ,則 為對稱矩陣。 - 冪等性(Idempotency):
若 ,則 為冪等矩陣。對於冪等矩陣,任意正整數次方均有 。 - 跡數與循環性質(Trace and Cyclic Property):
跡數具備循環置換不變性,即 。 - 冪等矩陣的性質:
- 冪等矩陣的特徵值(Eigenvalues)非 即 。
- 對稱冪等矩陣的秩(Rank)等於其跡數(Trace),即 。
- 若 的對稱冪等矩陣之 ,則存在值為 的特徵值,因此僅為半正定(Positive Semi-definite),而非正定(Positive Definite)。
解題方法
給定 為 階矩陣且 , 為 正定矩陣(可逆且 )。
定義矩陣 ,此即線性迴歸中的投影矩陣( 階)。
1. 驗證冪等性
中間相乘部分 ,故:
由此可知 為冪等矩陣,且對任意正整數 ,皆有 。
2. 計算跡數(Trace)
利用跡數的循環性質:
因 ,可得:
3. 計算矩陣的秩(Rank)
因 為對稱冪等矩陣,其秩等於其跡數:
第 8 題5 分
- Let be a sequence of independently and identically distributed random vectors with finite variances. Consider the following two simple regressions:
and
Let and be the LS estimators for and , respectively. Suppose that and are, respectively, consistent for and under the condition: . Which of the following is right?
(a) if .
(b) if , where is uncorrelated to , and .
(c) if .
(d) if and are two independent -distributed random variables.
(e) All of the above choices (a)-(d).
登入後即可作答並保存紀錄。
核心觀念
本題考查無截距項簡單迴歸(Simple Regression without Intercept)之最小自我乘法估計量(OLS / LS Estimator)的一致性(Consistency)、極限機率值(Probability Limit),以及**疊代期望值法則(Law of Iterated Expectations, LIE)**與卡方分配動差之應用。
-
無截距項迴歸之 LS 估計量的一致性:
對於模型 ,其 OLS 估計量為:根據大數法則(LLN),當樣本數 時, 一致收斂至:
-
真實資料生成過程(DGP)與外生性條件:
給定條件 ,真實條件期望值為 。透過疊代期望值法則,可推導出 與 之關聯。 -
自由度為 1 之卡方分配動差:
若 ,則其期望值為 ,變異數為 ,二階動差為:
解題方法
步驟一:推導 與 之理論關係式
由題意, 為真實模型且滿足嚴格線性外生條件 ,因此:
在模型 中, 之一致估計目標 為:
利用疊代期望值法則展開分子 :
將分子代回 的表示式中,可得關鍵通式:
選項分析
- (a) 正確:
若 ,將其代入通式: 故 (a) 敘述正確。
第 9 題5 分
- Assume that is a sequence of independently and identically distributed random vector. In particular, the random vector has the distribution , where is a matrix:
Suppose that is generated by the following regression:
for some and . Define two estimators: and . Which of the following is right?
(a) is consistent for if and .
(b) is consistent for if and .
(c) is consistent for if and .
(d) , as , if , and .
(e) None of the above choices (a)-(d).
登入後即可作答並保存紀錄。
核心觀念
本題探討計量經濟學中估計量的一致性(Consistency)與漸近變異數(Asymptotic Variance),核心理論基礎包括:
- 弱大數法則(Weak Law of Large Numbers, WLLN)與機率收斂(Convergence in Probability):
若隨機樣本列為 i.i.d.,其樣本平均數機率收斂至母體期望值: - 多變量常態分配之共變異數結構:
隨機向量 ,且期望值皆為零,其二階動差滿足: - 多元常態分配高階動差(Isserlis' Theorem / Wick's Theorem):
若四個隨機變數共同服從平均數為零的常態分配,則四階動差可拆解為二階共變異數之乘積和:
解題方法
將真實模型 代入兩個估計量中進行展開:
1. 估計量 的極限行為
根據大數法則(WLLN):
因此,由連續對應定理(Continuous Mapping Theorem):
欲使 為 的一致估計量(即 ),對任意 恆成立,充要條件為:
2. 估計量 的極限行為
根據大數法則(WLLN):
因此:
欲使 為 的一致估計量,充要條件為:
3. 差值估計量之漸近變異數
考慮兩估計量的差值:
定義隨機變數 。由於樣本為獨立同分配(i.i.d.),其乘上 後的變異數為:
第 10 題5 分
- Let be a sequence of independently and identically distributed random vectors, where has the logistic probability density function:
for and for some , is a -distributed random variable which is independent of , and for all 's. Consider a simple regression:
where is the regression coefficient, and is the regression error. Let be the least squares estimator for . Which of the following is right?
(a) converges in probability to , as .
(b) converges in probability to , as .
(c) converges in probability to , as .
(d) converges in probability to , as .
(e) None of the above choices (a)-(d).
登入後即可作答並保存紀錄。
核心觀念
本題考查計量經濟學與數理統計中的三大核心主題:
- 無常數項簡單迴歸(Regression Through the Origin)的最小平方法估計量(OLS Estimator):
對於無截距項模型 ,最小平方法估計量公式為: - 大數法則(Law of Large Numbers, LLN)與連續對映定理(Continuous Mapping Theorem, CMT):
根據辛欽弱大數法則(Khinchin's WLLN),樣本平均數依機率收斂至母體期望值: 再由連續對映定理,若分母極限非零,則: - 獨立隨機變數期望值的性質與常用分配動差:
- 乘積期望值:若 與 獨立,則 。
- 卡方分配動差:若 ,其期望值為自由度 。
- Logistic 分配動差:題目給定 為對稱於 0 的尺度參數 之 Logistic 分配,其變異數為 。
解題方法與詳細推導
步驟一:寫出 OLS 估計量 之表示式
考慮無常數項迴歸模型 ,目標為最小化殘差平方和 。一階條件為:
步驟二:代入真實資料生成過程(DGP)
根據題意,資料生成關係為 。代入 的分子:
將分子與分母同除以 :
步驟三:利用大數法則探討機率極限(Probability Limit)
因為 為獨立同分配(i.i.d.),各動差存在且有限:
- 分母極限:
Part II. 統計個案分析(50%,選擇題,答案可能不只一個選項,每題 5 分)
Ash Ketchum, the owner of Pikachu, surveyed attack powers of all 23 Pokémon monsters after he won the 2022 Pokémon World Championships. He studied all monsters for five years and regress attack powers of Pokémon on their health points (HPs). A simple linear regression analysis was specified as
where and indicate subscripts for monster and year, respectively. is the attack power, and is the HP. Regression was performed by ordinary least squares (OLS) method with results below.
Summary
| R-sq | 0.090 |
| Adj. R-sq | ▢(空白) |
| SE | 56.800 |
| N | 115.000 |
ANOVA
| DF | SS | MS | F | P-value | |
|---|---|---|---|---|---|
| Regression | 1.000 | 36129.233 | 36129.233 | 11.199 | 0.001 |
| Residual | 113.000 | 364560.781 | 3226.202 | ||
| Total | 114.000 | 400690.014 |
Regression estimates
| Coefficient | SE | t-value | P-value | |
|---|---|---|---|---|
| Intercept | 91.344 | 17.096 | 5.343 | 0.000 |
| HP | 1.732 | 0.518 | 3.346 | 0.001 |
第 11 題5 分
- The data form of Ash's study could be known as
(a) Cross-sectional
(b) Time-series
(c) Panel
(d) Count
(e) Categorical
登入後即可作答並保存紀錄。
核心觀念
本題評量計量經濟學與統計學中**資料型態(Types of Data)**的分類與辨識。實證資料依據「個體維度(Cross-section dimension)」與「時間維度(Time dimension)」的組合,主要分為以下三種基本架構:
- 橫斷面資料(Cross-sectional Data):在單一時間點上,蒐集多個不同個體(如個人、公司、國家、寶可夢)的資料。資料標註通常只有個體指標 ()。
- 時間序列資料(Time-series Data):針對單一特定個體,在連續多個不同時間點進行重複觀察。資料標註通常只有時間指標 ()。
- 縱橫資料/面板資料(Panel Data / Longitudinal Data):結合橫斷面與時間序列,在多個不同時間點()對同一群固定的多個個體()進行持續追蹤與重複測量。資料標註具有雙重下標 ,總樣本數為 (平衡面板資料,Balanced Panel Data)。
解題方法
- 辨識研究設計的觀測維度:
- 題目明確指出研究對象為 隻寶可夢(),代表個體維度,以指標 表示。
- 研究持續追蹤了 年(),代表時間維度,以指標 表示。
- 迴歸方程式設定為: 其中下標 與 分別對應寶可夢與年份。
- 驗算總樣本數:
- 摘要表中的樣本數為 ,恰好等於 筆觀測值,符合平衡面板資料(Balanced Panel Data)之特徵。
- 結論:
同時具備個體維度 與跨期追蹤時間維度 的二維資料型態,即為面板資料(Panel Data,亦稱縱橫資料或追蹤資料)。
第 12 題5 分
- Ash wants to know the distribution of the variable, i.e., the attack power, by QQ-plot. To plot the QQ-plot, he obtained a figure in the right. Which one(s) of following statement should be true?
🖼️【此處有附圖,請對照原卷:normalized y 的 QQ-plot,橫軸 −3~2】
(a) Data are not randomly sampled
(b) Data are left skewed
(c) Data are right skewed
(d) Data are fat-tailed
(e) Data are generally a Poisson distribution
登入後即可作答並保存紀錄。
核心觀念
本題的核心觀念為常態分位圖(Normal Quantile-Quantile Plot,簡稱 Normal Q-Q Plot)的分佈特徵判讀。
Normal Q-Q Plot 是將樣本分位數(Sample Quantiles,縱軸)與標準常態分佈的理論分位數(Theoretical Quantiles,橫軸)進行對照的圖形工具:
- 常態性(Normality):若資料服從常態分佈,散布點會大致落在對角參考直線( 線)上。
- 偏態(Skewness):
- 右偏(正偏,Right-skewed):右側存在較長尾部(極端大值顯著大於常態預期),左側被下界限縮。在 Q-Q Plot 上,左端點高於參考線、右端點亦高於參考線,整體呈現開口向上的 U 型(Convex / 向上彎曲)。
- 左偏(負偏,Left-skewed):左側存在較長尾部(極端小值顯著小於常態預期)。在 Q-Q Plot 上,整體呈現開口向下的倒 U 型(Concave / 向下彎曲)。
- 峰態與尾部(Kurtosis / Tail Heaviness):
- 厚尾(Fat-tailed / Heavy-tailed):兩端極端值皆比常態分佈更極端,左側散布點低於參考線、右側高於參考線,呈現典型 S 型。
- 薄尾(Thin-tailed / Light-tailed):左側散布點高於參考線、右側低於參考線,呈現反 S 型。
- 抽樣獨立性:Q-Q Plot 僅反映邊際經驗分配的幾何形狀,無法用以診斷樣本點之間是否具備獨立隨機抽樣性質。
解題方法
- 確認坐標軸設定:
圖形橫軸為標準常態理論分位數(範圍約 至 ),縱軸為標準化後的攻擊力 樣本分位數(Normalized )。 - 觀察點列相對於參考直線的偏離趨勢:
- 在左側(理論分位數為負值處),樣本分位數高於參考直線(即負向極端值沒有常態分佈預期的那麼負,左側被截斷或擠壓)。
- 在右側(理論分位數為正值處),樣本分位數隨著理論值增加而急遽向上攀升,且明顯位於參考直線上方(右側有極端大值,長尾延伸)。
- 整體散布點呈現明顯的 U 型(Convex)彎曲趨勢。
- 對應分配型態:
點列整體向上彎曲(U 型)即為統計學上標準常態 Q-Q Plot 代表**右偏(Right-skewed)**的經典特徵。
第 13 題5 分
- Based on information from ANOVA, the unconditional variance of , which is equal to , is about
(a) 56.800
(b) 61.523
(c) 3226.201
(d) 3514.825
(e) 4058.211
登入後即可作答並保存紀錄。
核心觀念
本題的核心觀念為變異數分析表(ANOVA Table)中的平方和分解以及**被解釋變數的無條件變異數(Unconditional Variance)**之定義。
- 總平方和(Total Sum of Squares, ):
衡量被解釋變數 圍繞其樣本平均數 的總離差平方和: - 無條件樣本變異數(Unconditional Sample Variance of ):
在不考慮任何自變數 資訊的前提下, 自身的樣本變異數定義為: 其中 為總自由度, 為總均方(Total Mean Square)。 - 條件變異數(Conditional Variance of given ):
給定自變數 後, 圍繞迴歸線的離散程度,即殘差變異數(Residual Variance),由殘差均方()估計:
解題方法
題目要求計算 的無條件變異數(Unconditional Variance),並已明確給出公式為:
由 ANOVA 表中可直接讀取對應數據:
- 總離差平方和
- 總自由度 (因總樣本數 ,故 )
代入公式計算:
選項分析
第 14 題5 分
- Ash does not interpret the volatility of by standard error but by variance. Which of following one(s) is right statement(s) of the standard error (SE) under assumptions of i.i.d.?
(a) SE is a non-linear function of
(b) SE is a biased estimator
(c) SE is an inconsistent estimator
(d) SE has the same unit of
(e) SE is more volatile than the variance
登入後即可作答並保存紀錄。
核心觀念
本題探討在古典線性迴歸模型(或樣本資料下)中,殘差標準誤(Standard Error of Regression, 或 )與殘差變異數( 或 )的統計性質比較,包含:
- 非線性轉換(Non-linear function):標準誤 ,透過開平方根與平方和運算,為應變數 的非線性函數。
- 估計量的偏誤性(Bias)與詹森不等式(Jensen's Inequality):無偏估計量經過非凹/非凸轉換後性質會改變。雖然 是母體變異數 的不偏估計量(),但因平方根函數 為嚴格凹函數(strictly concave),由詹森不等式可知 ,故 為 的有偏估計量(向下偏誤,underestimated)。
- 一致性(Consistency):由連續對映定理(Continuous Mapping Theorem, CMT)或大數法則,若 ,則 ,因此 是 的一致估計量(Consistent estimator)。
- 測量單位(Unit of measurement):變異數之單位為原變數單位的平方,而標準誤透過開根號還原,其單位與 相同。
解題方法
根據迴歸摘要報表,題幹中的 SE = 56.800 即為迴歸估計之標準誤(Standard Error of the Regression / Root MSE):
我們針對標準誤 之數學定義、抽樣分配特性(偏誤性、一致性)、單位以及與變異數 之波動度逐項驗證。
選項分析
-
(a) SE is a non-linear function of 【正確】
迴歸標準誤定義為:其中包含對應變數 的平方和以及整體開平方根運算,顯然不具備線性疊加性質(),因此 是 的非線性函數。
-
(b) SE is a biased estimator【正確】
在常態與 i.i.d. 假設下,,因此 (樣本變異數或 MSE 是母體變異數 的不偏估計量)。
然而,平方根函數 在 為嚴格凹函數()。根據詹森不等式(Jensen's Inequality):
第 15 題5 分
- In the ANOVA table, MS indicates
(a) Mean of Squares
(b) Sum of Squares divided by corresponding degree of freedom
(c) Mode of Symbol
(d) Max Security
(e) Maximum of Squares
登入後即可作答並保存紀錄。
核心觀念
在變異數分析表(ANOVA table)與線性迴歸分析中,MS 代表 Mean Square(均方,或稱 Mean of Squares)。
其統計學上的核心定義為:將平方和(Sum of Squares, SS)除以其對應的自由度(Degrees of Freedom, DF):
具體在單變量線性迴歸中:
- 迴歸均方(Mean Square of Regression, MSR):
- 殘差均方/誤差均方(Mean Square of Error / Residual, MSE):
亦為迴歸模型誤差項變異數 的不偏估計量()。
解題方法
本題屬於觀念與定義題。題幹特別註明「答案可能不只一個選項」(多選題形式)。
判斷步驟如下:
- 檢視縮寫意涵:在 ANOVA 表中,縮寫 MS 即為 Mean Square(均方,亦可表述為 Mean of Squares)。
- 檢視計算定義:檢視 ANOVA 表各欄位關係,數值欄位由左至右分別為 、、、。由表可知:
例如:
- 迴歸列:
第 16 題5 分
- The adj. R-sq is not presented in the tables. The adj. R-sq is equal to.
(a) 0.082
(b) 0.090
(c) 0.910
(d) 0.918
(e) 0.995
登入後即可作答並保存紀錄。
核心觀念
本題評量多元/簡單線性迴歸分析中「調整後判定係數」(Adjusted Coefficient of Determination, 或 )的定義與計算。
-
未調整判定係數():
表示應變數的總變異中,可由迴歸模型解釋的比例。但 具有「只要模型中增加自變數,其值絕不減少」的缺點。
-
調整後判定係數( 或 ):
為了懲罰模型中過多無效的自變數,調整後判定係數考慮了自由度的損失,以不偏估計量形式的均方(Mean Square)取代平方和:其中:
- 為樣本數(Total DF )。
- 為自變數個數(此處僅有自變數 HP,故 )。
- 誤差自由度 。
- 總自由度 。
-
與 之轉換關係式:
解題方法
本題有兩種計算途徑,皆可精確求得答案:
方法一:由變異數分析表(ANOVA Table)直接計算
從 ANOVA 表中可提取各項數值:
- 殘差均方
- 總平方和 ;總自由度
- 總均方 為:
代入調整後判定係數公式:
第 17 題5 分
- What is SS of the independent variable ?
(a) 106.112
(b) 3226.202
(c) 10240.904
(d) 400690.014
(e) 36129.232
登入後即可作答並保存紀錄。
核心觀念
本題評量簡單線性迴歸模型(Simple Linear Regression, SLR)中,**自變數離差平方和()與斜率估計式標準誤()**之間的代數關聯。
在古典簡單線性迴歸模型 中:
- 自變數的離差平方和(Sum of Squares of ):
- 隨機誤差變異數 的不偏估計量為殘差均方(Mean Squared Error, ):
- 斜率最小平方法估計量 的抽樣變異數與標準誤公式為:
- 迴歸平方和()與斜率估計式的關係為:
解題方法
題目要求計算自變數 的離差平方和,即 。由報表已知數據,有兩種標準解法:
方法一:利用斜率估計式的標準誤()求解
由迴歸係數估計表可知:
- 斜率估計值
- 斜率標準誤
由 ANOVA 表可知殘差均方:
根據公式 ,移項整理得:
代入數值:
方法二:利用迴歸平方和()與斜率估計值求解
由 ANOVA 表與迴歸係數估計表可知:
- 迴歸平方和
- 斜率估計值 (亦即 )
根據公式 ,移項整理得:
代入數值:
報表內部一致性比對與選項還原
注意到上述兩種標準計算的結果落在約 之間,但檢視題目選項:
- (c) 選項為
- 若以原始報表未四捨五入的精確值回推:
第 18 題5 分
- Given the data form (preferring to the answer in question 1 of this case study), which of following statement(s) is(are) true?
(a) The cross-sectional variations in the data generally violate the normality assumption in a large sample
(b) The time series variations in the data generally violate the independence assumption
(c) We can further specify a moving-average regression model for the residual to fit the identical assumption.
(d) The cross-sectional variations in the data generally can be modeled by an autoregressive regression model for the residual .
(e) Regression model using panel form of data cannot be estimated by maximum likelihood method.
登入後即可作答並保存紀錄。
核心觀念
本題的核心在於縱橫資料(Panel Data / Longitudinal Data)的基本特性與古典線性迴歸假說的檢驗:
- 資料型態(Panel Data):本研究追蹤了 隻寶可夢在 年間的攻擊力與血量,同時具有橫截面(個體 )與時間序列(時間 )維度,屬於平衡縱橫資料(Balanced Panel Data)。
- 古典線性迴歸的誤差項假說(Classical OLS Assumptions):
- 獨立性(Independence / No Autocorrelation):(當 或 )。
- 常態性(Normality):。
- 同質變異數(Homoskedasticity / Identical Distribution):。
- 時間維度的自相關性(Serial Correlation):同一觀察個體在不同時間點的衝擊往往存在持續性(Persistence),導致時間序列變異(Time series variations)通常違反誤差項相互獨立的假說。
解題方法
從題幹給定之資料架構切入:
- 分析時間維度的相依結構:同一隻寶可夢第 年的殘差 與第 年的殘差 很可能存在序列相關(Autocorrelation),即 (),這直接違反了殘差獨立性假說(Independence assumption)。
- 檢驗時間序列模型(AR / MA)的適用對象:自迴歸(AR)與移動平均(MA)模型建立在具有明確時間前後順序的維度上,用以捕捉序列自相關(解決獨立性違反問題),而非橫截面維度,亦非用來修正「同分布(Identical / Homoskedasticity)」假說。
- 檢驗估計方法的通用性:縱橫資料模型(如隨機效果模型 Random Effects Model)常假設個體隨機效應與誤差項服從常態分配,可直接以最大概似估計法(Maximum Likelihood Estimation, MLE)進行估計。
選項分析
第 19 題5 分
- Ash considered an alternative regression model that include a new variable and year fixed effect:
indicates year fixed effect. New variable, , is the time trend, which is equal to 1, 2, 3, 4 and 5 for year 2018, 2019, 2020, 2021 and 2022, respectively. , , … and are dummies for year fixed effects. Which one(s) of following statements is(are) true?
(a) Coefficient is generally negative.
(b) Average of to is zero.
(c) Results can be estimated by OLS
(d) A linear combination of , , , … and could be a constant.
(e) There is a perfect collinearity issue.
登入後即可作答並保存紀錄。
核心觀念
本題旨在測驗多元迴歸模型中的虛擬變數陷阱(Dummy Variable Trap)與完全共線性(Perfect Multicollinearity):
- 完全共線性定義:當模型中的解釋變數(含常數項 )存在不全為零的常數 ,使得 恆成立時,資料矩陣 的行向量線性相依,導致 不可逆(),OLS 估計式 無法唯一求解。
- 虛擬變數陷阱:若一類別變數有 個互斥且穷盡的類別(此題為 5 個年度),其虛擬變數總和恆等於常數項()。若模型同時包含截距項與所有 個虛擬變數,就會陷入虛擬變數陷阱。
- 時間趨勢與年度虛擬變數的代數關係:時間趨勢變數 本身即為年度虛擬變數的線性組合: 若再計入截距項與五個年度虛擬變數,模型中同時存在兩組獨立的完全共線性關係。
解題方法
檢視設定的模型:
- 常數項與虛擬變數的共線性:
因為資料涵蓋 2018 至 2022 共 5 個年度,每筆觀察值必且僅屬於其中一個年度,因此對所有觀察值恆有: 這直接等於模型中的常數項對應之行向量(全為 1 的常數),此即經典的虛擬變數陷阱。 - 時間趨勢與虛擬變數的共線性:
時間趨勢項定義為 ,因此: 這表示 完全可由年度虛擬變數線性表出。 - 推論結論:
- 變數之間存在線性組合等於常數的關係。
第 20 題5 分
- Ash also plans to estimate following regression model:
where and indicate subscripts for monster and year, respectively. is the attack power, and is the HP. Regression was performed by ordinary least squares (OLS) method. Which one(s) of following statements is(are) true?
(a)
(b)
(c)
(d)
(e) R-sq of regressing on is 0.073
登入後即可作答並保存紀錄。
核心觀念
本題評量**簡單線性迴歸(Simple Linear Regression, SLR)**中,將應變數()與自變數()互換時的估計量關係與判定係數性質:
-
迴歸斜率與相關係數的關係:
在 對 的迴歸模型 中,OLS 斜率估計量為:在 對 的逆向迴歸模型 中,OLS 斜率估計量為:
兩者之乘積恰為 Pearson 樣本相關係數的平方(即判定係數 ):
因此逆向迴歸的斜率估計量為:
-
判定係數之對稱性:
在簡單線性迴歸中,判定係數 等於兩變數樣本相關係數的平方 。由於 ,故 對 迴歸之 與 對 迴歸之 完全相同。
解題方法
由題幹給定的迴歸估計表與 ANOVA 表可取得下列數值:
- 對 的迴歸係數:
- 判定係數:(由 ANOVA 表亦可計算 )
- 計算逆向迴歸斜率 :
利用 ,代入數值: