115 年 國立臺灣大學財務金融研究所乙組《統計學(I)》
第 1 題
Ash Ketchum collected 100 samples of Pikachus and measured their weights. Few outliers are found, and which of following way(s) can be potential measures to reduce effect of outliers?
(A) Winsorize outliers.
(B) Remove outliers.
(C) Take natural logarithm of outliers.
(D) Take exponential of outliers.
登入後即可作答並保存紀錄。
本題主要在探討如何處理資料中的離群值 (outliers) 以減少其對分析結果的影響。
離群值是指資料集中與其他數據點顯著不同的數值。它們可能是測量錯誤、輸入錯誤,或是真實但極端的數據。若不妥善處理,離群值可能嚴重扭曲統計量(如平均數、標準差)和模型參數的估計。
處理離群值的方法主要有以下幾種:
- 移除離群值 (Remove outliers):直接將離群值從資料集中刪除。這是最直接的方法,但可能損失有用的資訊,且若離群值是真實的極端值,則移除會導致結果偏差。
- 溫索化 (Winsorize outliers):將離群值「縮小」到某個預設的百分位數上。例如,將低於 5% 分位數的數值改為 5% 分位數的值,將高於 95% 分位數的數值改為 95% 分位數的值。這種方法保留了資料點的數量,但減弱了離群值的極端影響。
- 轉換資料 (Transform data):對整個資料集或部分資料進行數學轉換,例如取對數、開根號等。這些轉換可以縮小資料的變異範圍,有時能將原本的離群值拉近到其他數據點的範圍內。例如,取對數轉換(如自然對數)常對偏態(skewed)的資料有效,能夠減少極大值的影響。
- 使用穩健統計方法 (Robust statistical methods):選擇對離群值不敏感的統計方法,例如使用中位數代替平均數,或使用中位數絕對偏差 (Median Absolute Deviation, MAD) 代替標準差。
第 2 題
Ash Ketchum studies the power of four different PocketMons, and four independent groups are observed, each with the same sample size. The sample means and pooled mean square error from a one-way ANOVA are given below:
| Group | Name of PockMon | Sample Mean |
|---|---|---|
| 1 | Purin | 70 |
| 2 | Pikuchu | 75 |
| 3 | Zenigame | 78 |
| 4 | Fushigidane | 82 |
| Each group has sample size n=10. The pooled mean square error is . The researcher is interested in testing whether the average power of Pikuchu, Zenigame, and Fushigidane differs from the Purin. Ash Ketchum decides to use contrast test for this hypothesis testing, where ; with the contrast . | ||
| Which of following one(s) is correct? | ||
| (A) The contrast coefficients must sum to zero. | ||
| (B) The test statistic follows a t-distribution with 39 degrees of freedom. | ||
| (C) The t-value is about -3.80. | ||
| (D) A significant contrast implies that the overall ANOVA F-test is significant. |
登入後即可作答並保存紀錄。
核心觀念
對 組平均數的線性組合
若係數滿足 ,則稱為對比(contrast)。本題的假設為
因此對比係數為 、,係數總和為 。
在單因子變異數分析中,對比檢定使用組內均方誤差 估計共同變異數,檢定統計量為
其自由度為 ANOVA 的組內誤差自由度 。
解題方法
先用樣本平均數估計對比值:
四組樣本數皆為 ,所以對比估計量的標準誤為
第 3 題
Let , where . Consider the event . Which of the following value(s) can the probability take for some value of ?
(A) 0
(B) p
(C) 1-p
(D) 0.999
登入後即可作答並保存紀錄。
本題考驗對 Bernoulli 分佈的理解,以及事件機率的計算與範圍。
核心觀念:
- Bernoulli 分佈:一個隨機變數 服從 Bernoulli() 分佈,表示 只有兩個可能的取值:1(成功)與 0(失敗)。其機率質量函數為 且 。
- 事件的機率:計算特定事件發生的機率。
- 機率的範圍:任何事件的機率都必須介於 0 和 1 之間(包含 0 和 1)。
題目設定:
- , 。
- 事件 。
- 要求 可能的取值。
分析事件 A:
由於 只能取 0 或 1,我們需要分別討論這兩種情況:
情況 1:
此時,。
。
事件 A 的條件變為 。
由於 ,所以 。我們可以同時除以 :
。
所以,如果 且 ,則事件 A 成立。
此情況發生的機率是 ,其中 是指示函數。
這個機率是 如果 ,否則為 0。
情況 2:
此時,。
。
事件 A 的條件變為 。
由於 ,所以 。我們可以同時除以 :
。
所以,如果 且 ,則事件 A 成立。
此情況發生的機率是 。
這個機率是 如果 ,否則為 0。
計算 :
第 4 題
Let be a continuous random variable with probability density function:
, .
Define a new random variable . Which of the following statement(s) are correct?
(A) Y takes values in (0,1).
(B) The cumulative distribution of Y is , .
(C) The probability density function of Y is , .
(D) .
登入後即可作答並保存紀錄。
本題考驗對連續隨機變數的變數轉換,包括求新隨機變數的取值範圍、累積分佈函數 (CDF) 和機率密度函數 (PDF),以及期望值。
核心觀念:
- 隨機變數轉換:若已知隨機變數 的分佈,求由 轉換而來的新隨機變數 的分佈。
- 累積分佈函數 (CDF):。
- 機率密度函數 (PDF):。
- 期望值 (Expectation): 或 。
題目數據與設定:
- 的 PDF: , for 。
- 新隨機變數 。
步驟一:確定 Y 的取值範圍
因為 的取值範圍是 ,且 是一個單調遞增函數在 上,所以:
當 時,。
當 時,。
因此, 的取值範圍是 。
選項 (A) 說 Y 取值在 (0,1),這是開區間,不包含端點 0 和 1。實際上,Y 可以取到 0 和 1。因此,選項 (A) 是錯誤的。
步驟二:求 Y 的累積分佈函數 (CDF)
由於 ,所以 。
對於 :
。
因為 的取值範圍是 ,且 ,所以我們只需要考慮 的部分。
。
第 5 題
Which of following estimator(s) could be biased to population mean given i.i.d. random variables with finite population mean ?
(A)
(B)
(C)
(D)
登入後即可作答並保存紀錄。
核心觀念
估計量 對母體平均數 的偏誤定義為
若 ,則 是不偏估計量;若在題目允許的情況下 ,則 會有偏。由於 為獨立同分配,對每個 都有 。
解題方法
逐一計算選項的期望,並與 比較。題目中的(B)寫成 ,未標示加總或平均;以下依照乘積解讀。由獨立性,,只要平均數有限,乘積的期望就存在。
選項分析
-
(A):不會有偏。
因此 。
-
(B):可能有偏。
由獨立性可得
因而偏誤為 ,只有在 時才為零,一般不等於零,所以此選項可能有偏。
-
(C):不會有偏。
第 6 題
Let be i.i.d. random variables such that , where is known. Consider the hypothesis test versus . Denote is the quintile of the standard normal distribution. Which of the following is the power function of this test?
(A)
(B)
(C)
(D)
登入後即可作答並保存紀錄。
核心觀念
檢定 對 是右尾檢定。已知母體變異數 ,樣本平均數的標準化統計量為
在虛無假設成立時,,因此顯著水準 的拒絕域為 。
檢定力函數是母數取真值 時,拒絕虛無假設的機率:
解題方法
在母體平均數為 時,樣本平均數滿足 ,所以
由於拒絕域為 ,檢定力為
第 7 題
Let . Denote sample average , , and . Then
(A)
(B) is a not a BLUE.
(C) is an unbiased estimator of .
(D) is a consistent estimator of .
登入後即可作答並保存紀錄。
核心觀念
- 樣本平均與期望:若 ,則
- 無偏性 (Unbiasedness):估計量 為 無偏 若 (此處 為母體參數 )。
- 一致性 (Consistency):估計量 一致 若 ,等價於 且 。
- BLUE(Best Linear Unbiased Estimator):在所有線性且無偏的統計量 ()中,變異數最小的即為 BLUE。對正態且同方差的樣本, 即為 的 BLUE。
解題方法
- 先對每一個定義的量寫出其期望與變異數。
- 判斷「常數」與「隨機變數」的本質差異。
- 依照無偏性、一致性及 BLUE 的條件逐一驗證選項。
各選項分析
| 選項 | 判斷 | 推導與說明 |
|---|---|---|
| (A) | 錯誤 | 為兩個觀測值的算術平均,屬於隨機變數。除非另有額外資訊(如 ),否則不會恆等於常數 。 |
第 8 題
Following previous question, please find Var().
(A) 0
(B)
(C)
(D) None of above
登入後即可作答並保存紀錄。
本題是接續上一題,要求計算 。
核心觀念:
- 變異數 (Variance):衡量隨機變數取值分散程度的指標。
- 統計量 (Statistic):一個基於樣本數據的量,其值不依賴於任何未知的母體參數。
題目設定:
- 承接上一題,,i.i.d.。
- 。
- 。
分析 :
正如上一題的詳細分析,。
根據期望值的性質,對於 i.i.d. 的 ,其期望值 。
所以,。
因此,。
現在我們需要計算 。
第 9 題
Let . Both and are unknown. Denote sample average , and . Then
(A) follows a Chi-square distribution.
(B) The chi-square distribution is used only when n is large enough.
(C) The confidence interval for is symmetric about .
(D) Increasing the sample size n narrows the confidence interval.
登入後即可作答並保存紀錄。
核心觀念
母體為常態分配時,樣本變異數有精確的卡方分配性質:
其中自由度為 。這項結果不要求樣本數很大,但要求樣本來自常態母體。
以此樞紐量可建立母體變異數 的信賴區間。令 表示自由度 的卡方分配第 分位數,則 信賴區間為
解題方法
先辨認樣本變異數的標準分配結果,再逐項檢查卡方分配的適用條件、信賴區間形狀,以及樣本數增加對區間精確度的影響。
卡方分配只取決於自由度,取值範圍為非負數,且通常呈右偏。因此變異數的信賴區間通常不會以 為中心對稱。增加樣本數會增加自由度、提升估計精確度;若比較時固定觀察到的 與信賴水準,區間會變窄。
第 10 題
Specify two regressions estimated on the same dataset; eq(1): ; eq(2): . is non-stochastic, and . is a dummy with values zero or one. and are error terms, which satisfy the classical linear regression assumptions. Which of following statement(s) is correct?
(A) Eq(1) cannot be estimated because perfect collinearity.
(B) Mean Squared Error (MSE) of eq(1) is smaller than MSE of eq (2).
(C) must be ranged between and .
(D) OLS estimator of is equal to OLS estimator of .
登入後即可作答並保存紀錄。
核心觀念
-
交互作用(Interaction)模型:
- eq(1) 可寫成
其中 與 為兩個互動項,分別在 與 時代表 的斜率。 - 此模型等價於「在 的兩個狀態下, 之迴歸係數可不同」的分組迴歸。
- eq(1) 可寫成
-
單純斜率模型:
- eq(2) 為
假設整體樣本共用同一條斜率 ,不考慮 的差異。
- eq(2) 為
-
多重共線性(Perfect collinearity):
- 若兩個解釋變數在樣本中線性相關(完全可以用另一個表示),OLS 無法唯一估計。
- 在 eq(1) 中, 與 只在 為 1 或 0 時分別非零,且不會同時為非零,故不構成完全共線。
-
OLS 的最小均方誤(MSE):
- 若模型正確規格(包含所有必要的交互項),則其殘差平方和(RSS)必不大於省略交互項的較簡模型。
- 同樣樣本、相同 、,eq(1) 包含 eq(2) 可視為「巢狀」模型,故其 均方誤(MSE)不會較大,且在 時嚴格較小。
-
斜率的範圍:
- 在分組迴歸中,整體斜率 為加權平均:
簡化後得到
其中 為 為 1 時 的比例,(除非全部 同值)。在大量樣本且誤差平均為 0 時,第二項趨近於 0,故 為 與 的加權平均,必介於兩者之間。若樣本全為同一 值,則 或 , 會等於唯一的斜率,但仍不會超出範圍。
- 在分組迴歸中,整體斜率 為加權平均:
-
截距的相等性:
- 兩模型的截距分別為 (eq(1))與 (eq(2))。
- 在 eq(1) 中,當 時 ,因 且 為非隨機, 為 的條件期望 。
- 在 eq(2) 中,同理 。兩式在同一樣本、相同 、相同資料下,估計 與 的 OLS 解皆等於樣本平均 (因 )——因此兩者相等。
解題方法
-
檢查多重共線性:檢驗 與 是否同時為常數倍。
- 若 ,則 ;若 ,則 。
- 只有在所有觀測值 同值時,兩變數會同時為零,導致共線。但題目未限定 為全同,故一般情況下 不存在完美共線。
-
比較 MSE:利用巢狀迴歸的性質。
- eq(1) 包含 eq(2) 的所有變項,且模型較為完全。OLS 以最小化 RSS 為目標,加入額外自變項只能降低或不變 RSS。
- 兩模型使用同一樣本且同一 ,故 eq(1) 的 MSE ≤ eq(2) 的 MSE,且在 時嚴格
<。
第 11 題
Consider the structural equation , where is endogenous that . Suppose we have an instrumental variable , and estimate via instrumental variable regression model, where . Which of the following statements are correct?
(A) For to be a valid instrument, it must satisfy both relevance () and the exclusion restriction ().
(B) If the instrument is weak, estimator is approximately unbiased in small samples.
(C) With a weak instrument, estimator can have larger bias than the OLS estimator.
(D) Asymptotic bias of is .
登入後即可作答並保存紀錄。
核心觀念
工具變數 要能識別內生變數 對 的因果效果,需滿足兩個條件:
- 相關性(relevance):,工具變數必須與內生解釋變數相關。
- 外生性/排除限制(exogeneity / exclusion restriction):,工具變數不得與誤差項相關。
題目給定的 IV 估計量為
將結構方程 代入,可得母體關係:
因此,只要 ,
若工具變數外生,,IV 估計量便能識別 。若工具變數與 的關聯很弱, 接近零,估計量會對抽樣波動特別敏感,造成弱工具變數問題。
解題方法
先用結構方程推導 IV 估計量的母體偏誤:
第 12 題
Consider the linear regression model , where , but the variance of the error term may be heteroskedastic. Which of the following statements are correct?
(A) The Breusch-Pagan test is based on regressing the squared OLS residuals on the original regressors (or a subset of them).
(B) The White test allows for heteroskedasticity of unknown functional form and includes cross-product terms of the regressors.
(C) Under the null hypothesis of homoskedasticity, the Breusch-Pagan and White test statistics are asymptotically chi-square distributed.
(D) If heteroskedasticity is present, the OLS estimator becomes biased and inconsistent.
登入後即可作答並保存紀錄。
本題考驗對線性迴歸模型中異質變異數 (Heteroskedasticity) 的檢定方法及其對 OLS 估計量的影響。
核心觀念:
- 異質變異數 (Heteroskedasticity):誤差項的變異數不是常數,即 隨樣本 而變化。
- 同質變異數 (Homoskedasticity):誤差項的變異數是常數,即 。
- OLS 估計量性質:
- 在同質變異數下,OLS 是 BLUE (最佳線性無偏差估計量)。
- 在異質變異數下,OLS 估計量仍然是無偏差且一致的,但不再是 BLUE,且標準誤的計算需要調整(例如使用 Huber-White 標準誤)。
- 異質變異數檢定:Breusch-Pagan 檢定和 White 檢定是用來檢測是否存在異質變異數。
題目設定:
- 線性迴歸模型:。
- 條件: (OLS 的無偏差和一致性條件)。
- 誤差項變異數:可能存在異質變異數。
逐一分析選項:
(A) The Breusch-Pagan test is based on regressing the squared OLS residuals on the original regressors (or a subset of them).
Breusch-Pagan (BP) 檢定 的步驟是:
- 估計原模型 得到 OLS 殘差 。
- 計算殘差平方 。
- 迴歸 對原模型中的解釋變數 (或其子集)。
- 檢定這個輔助迴歸的 是否顯著。
所以,BP 檢定確實是基於將殘差平方對解釋變數進行迴歸。
此敘述是正確的。
**(B) The White test allows for heteroskedasticity of unknown functional form and includes cross-product terms of the regressors
第 13 題
Given a function
, ,
please determine which of following statement(s) is correct.
(A) can be a probability density function.
(B) .
(C) .
(D) Denote odds = for a given number , then .
登入後即可作答並保存紀錄。
核心觀念
- 本題涉及 Logistic 分布(位置參數 、尺度參數 )。其累積分布函數(CDF)為
而概率密度函數(PDF)即為 的導數
- 檢驗 PDF 必須滿足 非負 以及 全域積分為 1。
- 期望值 以及 中心矩(特別是奇次中心矩)可由分布的對稱性直接判斷。
- Odds 與 log‑odds(logit)在 Logistic 分布中有簡潔的閉式表達式:
解題方法
- 驗證 為 PDF
- 明顯成立,因分子與分母皆為正。
- 以變數替換 ,則
- 設 ,則 ,積分變為
- 因此 , 為合法 PDF。
- 求期望值
- 由對稱性: 以 為中心對稱,亦即 .
- 對稱分布的期望等於對稱中心,故 。
- 若欲直接計算:
第 14 題
Neneneko recently won the lottery. He plans to spend the money to buy stocks. Neneneko is now studying the property of a certain stock X. Below presents the data of monthly returns of X in the past 8 months () and the corresponding market returns () as well as the risk-free rates (). Please answer questions 14 to 16 using the information.
| Month | (%) | (%) | (%) |
|---|---|---|---|
| 1 | 6 | 3 | 1 |
| 2 | 5 | 3 | 1 |
| 3 | 1 | -2 | 1 |
| 4 | 0 | -2 | 1 |
| 5 | -3 | 1 | 1 |
| 6 | -2 | 1 | 1 |
| 7 | 8 | 5 | 1 |
| 8 | 13 | 5 | 1 |
| Neneneko is interested in the systematic risk of stock X and would like to estimate the following regression: | |||
| Assume that and that $E[e_t | r_{mt}, r_{ft}] = 0$. Which of the followings are correct? | ||
| (A) The ordinary least square estimate . | |||
| (B) The ordinary least square estimate . | |||
| (C) The maximum likelihood estimate . | |||
| (D) The standard error of is 0.513. |
登入後即可作答並保存紀錄。
本題是關於金融市場中資產定價模型 (如 CAPM) 的線性迴歸分析,考驗對 OLS 估計的計算以及標準誤的理解。
核心觀念:
- 資本資產定價模型 (CAPM):描述資產預期報酬與系統性風險之間關係的模型。其簡化形式為 。
- 超額報酬 (Excess Return):資產報酬率減去無風險利率,即 和 。
- 迴歸模型:將資產的超額報酬對市場的超額報酬進行迴歸。
。
其中 (Beta) 是衡量資產系統性風險的指標。 - 普通最小平方法 (OLS):用於估計迴歸模型的參數。
- 標準誤 (Standard Error):衡量迴歸係數估計量精確度的指標。
題目數據與設定:
- 迴歸模型:
其中 (股票 X 的超額報酬)
(市場的超額報酬)
是截距 (Alpha), 是斜率 (Beta)。 - 樣本數 (8 個月)。
- 假設 。
步驟一:計算超額報酬
| Month | (%) | (%) | (%) | (%) | (%) |
|---|---|---|---|---|---|
| 1 | 6 | 3 | 1 | 5 | 2 |
| 2 | 5 | 3 | 1 | 4 | 2 |
| 3 | 1 | -2 | 1 | 0 | -3 |
| 4 | 0 | -2 | 1 | -1 | -3 |
| 5 | -3 | 1 | 1 | -4 | 0 |
| 6 | -2 | 1 | 1 | -3 | 0 |
| 7 | 8 | 5 | 1 | 7 | 4 |
| 8 | 13 | 5 | 1 | 12 | 4 |
步驟二:計算 OLS 估計量 和
OLS 估計量的公式:
首先計算平均值:
接下來計算分子和分母:
:
: (1.25, 1.25, -3.75, -3.75, -0.75, -0.75, 3.25, 3.25)
: (2.5, 1.5, -2.5, -3.5, -6.5, -5.5, 4.5, 9.5)
| 1.25 | 2.5 | 3.125 |
第 15 題
Following the regression specification in the previous question, Neneneko is interested in understanding the properties of the ordinary least square estimator under different hypotheses for the error term.
Assume that ; and that . Which of the followings are correct?
A1: ;
A2: , where N stands for normal distribution;
A3: , is independent of .
With the assumption of , which of the followings are correct?
(A) The ordinary least square estimator is BUE under A1.
(B) The ordinary least square estimator is BLUE under A2.
(C) Under A3, the ordinary least square estimator is biased.
(D) Under A3, given the full sample variance-covariance matrix of the error terms, the generalized least square estimator is BLUE.
登入後即可作答並保存紀錄。
核心觀念
令迴歸模型為 ,且 滿秩。題目給定外生性條件 ,其作用是確保 OLS 估計量不偏;但要判斷 OLS 是否具有「最佳」性質,還必須看誤差項的變異數與共變異數結構。
OLS 估計量為
其條件變異數為
其中 。若 ,Gauss–Markov 定理保證 OLS 是 BLUE:最佳線性不偏估計量,也就是在所有線性不偏估計量中變異數最小。
BUE 通常指最佳不偏估計量,範圍比 BLUE 更廣;Gauss–Markov 定理只比較線性不偏估計量,不能單憑它推出 OLS 是 BUE。
以下將 A1、A2、A3 視為各自要檢視的誤差假設:A1 只提供平均數與變異數資訊;A2 提供常態分配資訊;A3 在相同變異數條件下另加誤差間獨立。
解題方法
先把「不偏」和「最佳」分開判斷:
- 外生性 保證 。
- BLUE 還需要誤差共變異數矩陣的結構足以支持 Gauss–Markov 定理;只知道每個誤差項各自的變異數,不代表不同誤差項之間不相關。
- 若誤差共變異數矩陣 已知且正定,GLS 估計量為
第 16 題
Ideally, Neneneko should perform the following regression:
Suppose that Neneneko is careless and mess up his code. He performs the following three regressions instead:
(1)
(2)
(3)
Which of the followings are correct?
(A) (This likely refers to )
(B)
(C)
(D)
登入後即可作答並保存紀錄。
核心觀念
本題測驗的是「迴歸模型變形」與「遺漏變數造成的偏誤」(omitted‑variable bias)。
- 理想模型:
- 若把等式左右同時加上 (假設 為常數或至少與其他變數無關),可得到 真實的結構方程
此式說明:若只把 放入迴歸,而把 忽略,截距會吸收 的期望值;若把 當作唯一解釋變數,截距會吸收 。
遺漏變數與已納入變數之間的相關性(本案例中 與 / 必然相關)會導致 係數偏誤,而偏誤的大小取決於被遺漏變數的平均值與與已納入變數的共變異。
解題方法
-
將每個錯誤迴歸寫成與 (★) 的關係
- (1)
- (2)
- (3)
-
使用期望運算(假設誤差 均滿足OLS 的零均值與與解釋變數不相關的條件)得到每個模型的截距在母體層面的期望:
| 模型 | 變換後的母體等式 | 截距的期望值 |
|---|---|---|
| (★) (理想) | ||
| (1) | ||
| (2) | ↔ | |
| (3) |
其中 為風險無關利率的期望值(題目暗示「1%」)。
- 比較各截距的關係
- 與 :。
- 與 :。只有在 時才會剛好等於 ,一般情況下不等於 。
- 與 :。除非 且 才會得到「+1%」的差距,亦非必然。
第 17 題
Suppose that a data generating process is as the following:
However, cannot be directly observed. There are two observable proxies for :
We know that , , is independent of , , and . and are also uncorrelated to , , , and each other.
We now construct an additional regressor:
For or , denote as the OLS coefficient when we regress on . Which of the followings are correct?
(A)
(B)
(C)
(D)
登入後即可作答並保存紀錄。
核心觀念
- 測量誤差 (Errors‑in‑Variables):當解釋變數 只能以含誤差的代理變數 觀測時,OLS 估計會產生衰減偏誤 (attenuation bias)。
- Plim 公式(大樣本極限)
因為 與所有測量誤差皆獨立,故 。
- 若代理變數為 classical measurement error(與真實 不相關且誤差獨立同分布),則
其中 為 的變異, 為對應測量誤差的變異。
解題步驟
- 計算 的變異與與 的共變
- 的 plim
- 的 plim
- 構造新變數
- 變異
- 與 的共變
第 18 題
The number of red balls and blue balls in a bag is unknown, but it is known that the proportion, p, of red is either , , or . A sample of size 5, drawn with replacement, yields the sequence red, blue, blue, red, and blue. The maximum likelihood estimate for p is:
(A)
(B)
(C)
(D)
登入後即可作答並保存紀錄。
好的,這題是關於最大概似估計 (Maximum Likelihood Estimation, MLE) 的題目,核心概念是利用觀測到的樣本資料,找出最有可能產生這些資料的參數值。我們需要根據給定的樣本和可能的參數值,計算出每個參數值下的樣本出現機率(概似函數),然後找出使概似函數最大的那個參數值。
解題過程:
這是一個二項分佈 (Binomial Distribution) 的問題,因為我們是從袋子中有放回地抽取樣本,每次抽取都是獨立的,且每次抽取只有兩種結果:抽到紅球或藍球。令 為抽到紅球的機率。
我們有三種可能的 值:,,或 。
抽取的樣本大小為 。
觀測到的樣本序列是:紅、藍、藍、紅、藍。
這表示我們抽到了 2 個紅球 (R) 和 3 個藍球 (B)。
對於一個大小為 的樣本,其中有 個成功(在這裡是抽到紅球),則其二項分佈的機率質量函數 (Probability Mass Function, PMF) 為:
在我們的例子中,,(紅球的數量)。所以,對於一個給定的 ,觀測到 2 個紅球和 3 個藍球的機率是:
現在,我們需要計算在每個可能的 值下,這個機率值,也就是計算概似函數 在這三個點的值。
-
當 時:
-
當 時:
第 19 題
In the two-variable model:
Suppose that , , , , , and , where , , and are the column vectors with typical elements , , and , respectively. Furthermore, , , and are the transpose of , , and , respectively. Assume .
Now suppose you would like to make out-of-sample predictions about the dependent variable for one hypothetical observation for some . We can observe that and . Please estimate the expected value and variance of using the formula . Which of the following are correct?
(A)
(B)
(C)
(D)
登入後即可作答並保存紀錄。
核心觀念
本題考查無截距多元線性迴歸的普通最小平方法(OLS)、迴歸係數估計量的期望與變異數,以及樣本外預測值的期望與變異數。
令 ,則
在誤差項平均數為 、且 可逆時,
對固定的新觀測自變數 ,題目定義的預測值為 ,因此
解題方法
先由題目給定的內積組成 與 :
其行列式為 ,故矩陣可逆,且
因此
由 ,係數估計量的期望等於真實係數;本題的點估計分別為 、。
接著估計誤差變異數。殘差平方和為
共有 筆觀測、估計 個係數,因此無偏誤差變異數估計量為
第 20 題
Let the random variable have probability function
where . Which of the following are correct?
(A)
(B)
(C)
(D) If , the distribution of would have fatter tails than a standard normal distribution.
登入後即可作答並保存紀錄。
核心觀念
本題考離散隨機變數的動差,以及偏態係數與峰態係數的定義:
- 平均數:
- 變異數:
- 偏態係數:
- 峰態係數:;標準常態的峰態係數為 。超額峰態為 ,常用來比較分布的峰態與尾部厚薄。
解題方法
先依機率函數列出三個可能取值,再計算各階動差。由於 和 的機率相同,分布以 為中心對稱,奇數階動差會互相抵銷。
平均數為
二階動差為
因此變異數為
因為 ,三階中央動差就是三階原始動差:
第 21 題
Huang, Jiang, Tu, and Zhou (2015, RFS) constructed a market sentiment index, , and compared it against the seminal Baker and Wurgler (2006, JF) sentiment index, . They showed that the monthly variable can negatively predict the market returns in the following month. Which of the following are correct?
🖼️【此處有附圖,請對照原卷】
(A) Suppose the AR(1) coefficient of is 0.98. Since it is smaller than 1, we do not have to worry about the unit root problem.
(B) If a predictor, such as or , has a unit root, regressing future market returns on it would definitely produce a spurious regression.
(C) If a time series has a unit root, taking a first difference always removes the unit root.
(D) If a time series has a unit root, an exogenous shock to the series in a given time period could have a permanent effect on all future realizations.
登入後即可作答並保存紀錄。
核心觀念
- 單根過程(Unit Root Process)與衝擊持久性(Persistence of Shocks):
若時間序列具有單根(即整合階數 ,如 的隨機漫步),外生衝擊(shock)將永久留在系統中。對任意落後期 ,衝擊反應函數滿足以:
因此衝擊具有永久性影響(permanent effect);相對地,定態序列(stationary, )的外生衝擊僅具暫時性影響(transitory effect),隨時間推移衰減至零。 - 近單根問題(Near-Unit-Root Problem)與持續性預測變數(Persistent Predictors):
在有限樣本中,若 AR(1) 係數接近 1(例如 0.98),屬於「近單根(near-unit-root / local-to-unity)」序列。- 若 0.98 為估計值,因 OLS 存在著名的有限樣本向下偏誤(Dickey-Fuller bias),不能僅因估計值小於 1 就排除單根,仍須進行正式單根檢定(如 ADF 檢定)。
- 在資產報酬預測迴歸中,高持續性預測變數會導致嚴重的 Stambaugh 偏誤(Stambaugh bias),使得檢定統計量嚴重偏誤,故絕不能忽視單根或近單根問題。
- 整合階數與差分去除單根(Order of Integration and Differencing):
一階差分 僅能消除一個單位根。若序列含有二階或更高階單根(),一階差分後仍為非定態的 序列,依然含有單根。 - 虛假迴歸(Spurious Regression)的形成條件:
經典虛假迴歸(Granger & Newbold, 1974)是指兩個彼此獨立的非定態 序列互相迴歸,造成殘差非定態、判定係數 與 統計量發散膨脹。若被解釋變數為定態市場報酬率 ,此迴歸為非平衡迴歸(unbalanced regression),大樣本下估計係數收斂至 0,並不「必然(definitely)」產生虛假迴歸。
解題方法
原卷附圖顯示 1965 年至 2010 年間的月資料市場情緒指數走勢,縱軸範圍約在 -2 至 3 之間,實線代表偏最小平方法情緒指標 ,虛線代表 Baker-Wurgler 情緒指標 ,灰色陰影區間代表經濟衰退期(如 NBER 衰退期);圖形反映兩指數走勢呈現高度重疊、緩慢均值回歸且自我相關極高之特徵。
本題評量時間序列計量經濟學中單根的統計性質、衝擊反應、有限樣本近單根推論,以及預測迴歸的計量問題。解題切入點為依據單根過程的數學定義與實證計量文獻,逐一檢驗各選項敘述的邏輯嚴謹性與例外條件。
選項分析
- 選項 (A) 錯誤:
- 若 0.98 為樣本估計值:在小樣本下,若真實資料生成過程具有單根(),OLS 估計量 具有顯著的向下偏誤(downward finite-sample bias)。因此估計值為 0.98 完全可能來自一個具有單根的母體,不能單憑數值小於 1 就斷定無單根問題,仍須進行單根檢定(如 ADF 檢定、Phillips-Perron 檢定)。
- 若 0.98 為真實母體係數: 屬於典型的近單根(near unit root / local-to-unity)過程。
Chen, Huang, Lin, and Sheng (2022, MS) studied how access to finance affects the business of small and medium enterprises (SBEs) using data from Alibaba. The left figure shows the probability of credit access for SBEs of different credit scores. The right figure shows the volatility of the SBEs’ sales value growth in the following quarter.
🖼️【此處有附圖,請對照原卷】
第 22 題
Which of the following are correct?
(A) The researchers show that access to finance may causally reduce sales growth volatility.
(B) The discontinuity on the left figure undermines the reliability of the researchers’ intention.
(C) The negative slope on the left half of the right figure suggests that higher credibility correlates with lower sales growth volatility.
(D) The research may be invalidated if SBEs can proactively alter their credit scores around the discontinuity.
登入後即可作答並保存紀錄。
核心觀念
本題考查迴歸不連續設計(Regression Discontinuity Design, RD)。研究者利用信用分數門檻附近的差異,估計取得融資對後續銷售成長波動的影響。
若信用分數跨過門檻時,融資取得機率出現跳躍,而後續結果變數也在同一門檻出現跳躍,且門檻附近沒有其他因素同時改變,也沒有企業操弄分數,便可將結果變數的跳躍解讀為融資的因果效果。左圖的融資機率跳幅未達百分之百,表示這是模糊迴歸不連續設計;概念上以結果變數的門檻跳幅除以融資機率的門檻跳幅,估計門檻附近取得融資的因果效果。
解題方法
圖中橫軸是信用分數,門檻約在 ;左圖的信用取得機率在門檻處向上跳躍,右圖的下一季銷售額成長波動則在同一門檻處向下跳躍。右圖門檻左側的線也呈負斜率,表示分數愈高,波動愈低的相關趨勢。
判斷選項時,要分清楚兩種證據:門檻處的跳躍可用來評估融資的因果效果;圖上的一般斜率則呈現信用分數與波動的相關趨勢,不能單獨證明因果。RD 的因果解讀還須仰賴門檻附近其他條件連續、企業無法精確操弄信用分數等識別假設。
選項分析
第 23 題
Hvidberg (2023, RFS) studied the relationship between college majors and financial behaviors years after graduation. The left figure below shows the probability for high school students to get admitted into their first-choice schools and majors across their high school GPAs. The right figure below shows the probability of a default event 10 years after college graduation for Law and Business majors across their high school GPAs. Which of the following are correct?
🖼️【此處有附圖,請對照原卷】
(A) Overall, high school GPAs are positively correlated with the likelihood of a default event years afterwards.
(B) The discontinuity on the left figure suggests that college admission is contingent on certain GPA thresholds.
(C) The research indicates that being admitted into the first-choice school and major could causally reduce the likelihood of default for certain relevant majors.
(D) The research design is redundant, as the negative slopes on the right figure already support the argument that studying business causally reduces the likelihood of default.
登入後即可作答並保存紀錄。
核心觀念
本題考的是迴歸不連續設計(Regression Discontinuity Design, RD)。當某個處置機率在明確門檻處跳升,而結果變數也在同一門檻處出現跳躍,且門檻附近其他影響因素平滑變化時,可用結果的跳躍識別門檻附近的局部因果效果。
解題方法
左圖的橫軸是相對於 GPA 門檻的距離,縱軸是獲得第一志願學校與科系錄取的機率;門檻處的錄取機率明顯跳升。右圖呈現法學與商學科系學生畢業十年後的違約機率;門檻附近,圖中商學科系的違約機率在門檻處向下跳躍,表示跨過門檻、較有機會進入第一志願後,特定科系學生的違約機率降低。
因此,判斷重點是區分「圖上的斜率」與「門檻處的跳躍」:斜率反映 GPA 與違約機率的關聯;門檻處的結果跳躍,才是 RD 用來識別局部因果效果的依據。此因果解讀以門檻附近可比、且 GPA 無法被精確操控等 RD 條件成立為前提。
選項分析
第 24 題
Butler and Cornaggia (2011, JFE) studied whether access to finance improves productivity. They exploited a U.S. energy policy in 2005 that generated an exogenous increase in demand for corn due to ethanol production, since corn is a key input in ethanol production. They compared changes in corn yields before and after the policy across counties with high access versus low access to bank finance. They then applied the same analysis to soybeans as a control crop. Which of the following statements are correct?
(A) Soybeans are used as a control group because soybeans are not an input in ethanol production and thus their demand was not affected by the policy.
(B) The key identifying variation comes from a triple difference: across time (before vs. after the ethanol demand boom), across crops (corn vs. soybeans), and across counties (high vs. low access to finance).
(C) In the absence of the ethanol policy, corn yields in high-finance and low-finance counties would have to evolve in the same way over time for the study to be credible.
(D) The study identifies whether access to finance affects how strongly corn productivity responds to an exogenous increase in demand, rather than the unconditional level of productivity.
登入後即可作答並保存紀錄。
核心觀念
本題考查三重差分法(difference-in-differences-in-differences, DDD)。研究利用乙醇需求增加作為玉米需求的外生衝擊,檢驗金融可及性較高的地區,是否能比金融可及性較低的地區更有效提升玉米單位面積產量。研究以大豆作為對照作物,並以作物單位面積產量衡量生產力。Butler 與 Cornaggia 論文
三個差分維度是:
- 時間:乙醇需求增加前與增加後。
- 作物:玉米與大豆。
- 金融可及性:高與低。
令 表示作物 、金融可及性組別 、時期 的平均產量,三重差分可寫為:
這個差值衡量:乙醇需求增加後,高、低金融可及性地區的玉米產量變化差距,扣除同一期間大豆產量變化差距後,還剩下多少差異。
解題方法
判斷選項時,先辨認研究的處理組、對照組與衝擊,再檢查選項是否準確描述 DDD 的識別條件。玉米是受到乙醇需求衝擊的作物;大豆是沒有相同直接需求衝擊的對照作物;高、低金融可及性則用來比較生產者回應衝擊的差異。
第 25 題
Levitt (2021, REStud) studied whether people make good choices when facing important life decisions, such as quitting a job or ending a relationship. He created an online platform where individuals who were undecided about a major choice could receive a randomly assigned recommendation to either make a change or maintain the status quo. He then surveyed participants to measure whether they followed the recommendation and how their well-being evolved six months later. Which of the following are correct?
(A) The random assignment is treated as an instrumental variable in this study.
(B) If any participant fails to follow the recommendation, the conclusions of the study are no longer valid.
(C) If some participants always do the opposite of what the random recommendation suggests, the causal interpretation of the study may no longer be reliable.
(D) Because the recommendation is randomly assigned, the results of the study automatically generalize to the entire population.
登入後即可作答並保存紀錄。
核心觀念
本題考查隨機分派、工具變數(instrumental variable, IV)、不遵從(noncompliance)、單調性,以及內部效度與外部效度的區別。
Levitt 的研究以隨機分派的擲幣結果作為工具變數,鼓勵參與者「改變」或「維持現狀」,再觀察他們實際採取的行動及後續幸福感。研究利用分派結果造成的行動差異,估計採取行動對結果的因果影響。Levitt 研究摘要
令 表示隨機分派的建議, 表示是否實際改變, 表示後續幸福感。以二元變數表示時,IV 的 Wald 估計量為:
此因果解釋需要工具變數與潛在結果獨立、工具變數能影響實際行動、工具變數不直接影響結果,以及單調性等條件。
解題方法