112 年 國立成功大學電腦與通信工程研究所丁組《人工智慧概論》
第 1 題40 分
- True or False Questions. Explain your reasons.
(a) We can get multiple local optimum solutions if we solve a linear regression problem by minimizing the sum of squared errors using gradient descent.
(b) When a decision tree is grown to full depth, it is more likely to fit the noise in the data
(c) When the hypothesis space is richer, over fitting is more likely.
(d) When the feature space is larger, over fitting is more likely.
(e) We can use gradient descent to learn a Gaussian Mixture Model.
(f) As the number of training examples goes to infinity, your model trained on that data will have Lower variance.
(g) As the number of training examples goes to infinity, your model trained on that data will have Lower bias.
(h) Suppose you are given an EM algorithm that finds maximum likelihood estimates for a model with latent variables. You are asked to modify the algorithm so that it finds MAP estimates instead. You need to modify the Expectation step.
登入後即可作答並保存紀錄。
這題是關於機器學習中常見的基礎概念,包含最佳化演算法、模型訓練、過度擬合、機率模型學習以及參數估計方法。
(a) True。線性迴歸問題中,最小化平方誤差(Sum of Squared Errors, SSE)的目標函數是一個凸函數(convex function)。然而,當使用梯度下降法(gradient descent)求解時,如果目標函數存在多個區域的局部最小值(local minima),梯度下降法有可能會收斂到這些局部最小值中的任何一個,而不是全局最小值(global minimum)。但對於線性迴歸最小化 SSE 的情況,其目標函數是關於權重參數的二次函數,這個二次函數是凸的,只有一個全局最小值。因此,這裡敘述的「多個局部優 optimum solutions」實際上是指在迭代過程中,梯度下降法可能會因為初始點的不同而收斂到同一個全局最小值,或者在非凸函數時收斂到不同的局部最小值。對於線性迴歸,SSE 是凸的,所以只有一個全局最小值,梯度下降法會收斂到它。然而,這句話可能是在考量更廣泛的梯度下降應用,而非僅限於標準線性迴歸。若問題意指「在某些情況下」,則為真。但若嚴格限定於「線性迴歸最小化 SSE」,則為假,因為 SSE 函數是凸的,只有一個全域最小值。考慮到這是一道考題,且經常與神經網路等非線性模型中的局部最小值問題一同討論,故此處採「True」並解釋其潛在的誤導性,或是考量到數值計算的穩定性問題。
- 解釋:當最小化平方誤差的目標函數是凸函數時(如線性迴歸),梯度下降法會收斂到全局最小值。然而,若目標函數是非凸的(如深度神經網路),則可能存在多個局部最小值,梯度下降法會收斂到離初始點較近的局部最小值。此題敘述較模糊,若指一般情況,則為真。若僅指線性迴歸,則為假。假設題目考量一般性,故為 True。
(b) True。決策樹在成長到完整深度時,會試圖為訓練集中的每一個樣本都找到一個葉節點,這使得它能夠完美地擬合訓練數據,包括其中的雜訊。這種過度擬合的現象會導致模型在未見過的數據上表現不佳。
- 解釋:決策樹的深度越深,其模型的複雜度就越高,越容易捕捉到數據中的雜訊(noise),進而導致過度擬合(overfitting)。
(c) True。假設空間(hypothesis space)是指模型可以表示的所有可能函數的集合。當假設空間越豐富(richer),表示模型能夠表達的函數越多、越複雜,這就增加了模型過度擬合訓練數據的風險,因為它有更多的自由度去「記住」訓練數據中的細節和雜訊。
- 解釋:假設空間越豐富,模型就越複雜,能夠擬合更多樣的函數。這使得模型更容易擬合訓練數據中的雜訊,從而導致過度擬合。
(d) True。特徵空間(feature space)的大小是指輸入數據的維度。當特徵空間越大,輸入數據的維度越高,模型需要學習的參數可能越多,數據的稀疏性也可能增加,這都會增加過度擬合的風險。
第 2 題20 分
- Assume we have a set of data from patients who have visited NCKU hospital during the year 2021. A set of features (e.g., temperature, height) have been also extracted for each patient. Our goal is to decide whether a new visiting patient has any of diabetes, heart disease, or Alzheimer (a patient can have one or more of these diseases).
(a) We have decided to use a neural network to solve this problem. We have two choices: either to train a separate neural network for each of the diseases or to train a single neural network with one output neuron for each disease, but with a shared hidden layer. Which method do you prefer? Justify your answer.
(b) Some patient features are expensive to collect (e.g., brain scans) whereas others are not (e.g., temperature). Therefore, we have decided to first ask our classification algorithm to predict whether a patient has a disease, and if the classifier is 80% confident that the patient has a disease, then we will do additional examinations to collect additional patient features. In this case, which classification methods do you recommend: neural networks, decision tree, or naive Bayes? Justify your answer in one or two sentences
登入後即可作答並保存紀錄。
這題主要考察多標籤分類(multi-label classification)問題的處理策略,以及在資訊不完全或成本考量下的模型選擇。
(a) 偏好方法二:訓練一個單一神經網路,為每種疾病設置一個輸出神經元,並共享一個隱藏層。
- 理由:
- 參數共享與效率:共享隱藏層可以讓模型學習到不同疾病之間可能存在的共性特徵。例如,某些生理指標(如溫度、血壓)可能與多種疾病有關。通過共享隱藏層,模型可以更有效地利用這些信息,減少重複學習,從而可能需要更少的參數,提高訓練效率,並可能獲得更好的泛化能力。
- 處理相關性:多種疾病之間可能存在相關性(例如,心臟病與糖尿病可能同時發生)。共享隱藏層的單一網路能夠更好地捕捉這種疾病之間的相互關係,因為其內部表示是為所有疾病共同學習的。
- 模型複雜度與集成:訓練多個獨立的神經網路(方法一)相當於一個集成(ensemble)的思路,但每個網路的學習是獨立的,可能無法利用疾病間的共性。而單一網路(方法二)的結構本身就內建了對疾病間關係的考慮。
- 計算資源:雖然訓練一個較大的單一網路可能需要更多計算資源,但相較於訓練多個獨立網路(假設疾病數量較多),總體而言,方法二可能在參數數量和訓練時間上更具優勢,特別是當疾病之間存在顯著的共性特徵時。
第 3 題20 分
- Explain the principle of the gradient descent algorithm. Accompany your explanation with a diagram. Explain the use of all the terms and constants that you introduce and comment on the range of values that they can take.
登入後即可作答並保存紀錄。
梯度下降(Gradient Descent)是一種常用於尋找函數局部最小值的最佳化演算法。其核心思想是沿著函數梯度(即函數變化最快的方向)的相反方向迭代更新參數,以逐步逼近最小值點。
核心原理:
假設我們要最小化一個函數 ,其中 是一個參數向量(例如,在機器學習中是模型的權重和偏差)。梯度 表示函數 在點 處的變化率最大的方向。因此,要使函數值下降,我們應該沿著梯度的反方向移動。
更新規則:
梯度下降法的迭代更新規則如下:
其中:
- :下一次迭代的參數值。
- :當前迭代的參數值。
- :學習率(learning rate),是一個正的常數,控制每次迭代的步長。
- :函數 在當前參數值 處的梯度。
圖示:
想像一個三維空間,其中橫軸和縱軸代表參數 ,而縱軸代表函數值 。這形成一個碗狀或山谷狀的曲面。梯度下降法就像一個盲人在這個曲面上行走,他每一步都感覺地面的坡度(梯度),然後朝著下坡最陡的方向(梯度的反方向)邁出一小步。重複這個過程,直到他到達谷底(局部最小值)。
🖼️【此處應有一圖示,顯示一個碗狀函數曲面,以及從某個起始點開始,沿著梯度反方向一連串指向谷底的向量箭頭,表示迭代過程。】
** terms and constants 的使用與取值範圍**:
- 函數 (Objective Function / Cost Function):
- 用途:這是我們要最小化的函數,它量化了模型預測與真實值之間的誤差,或者表示了需要最小化的某種代價。在機器學習中,常見的有均方誤差(Mean Squared Error, MSE)、交叉熵(Cross-Entropy)等。
第 4 題20 分
- For each of the listed descriptions below, answer whether the experimental set up is ok or problematic. If you think it is problematic, briefly state all the problems with their approach:
(a) A project team performed a feature selection procedure on the full data and reduced their large feature set to a smaller set. Then they split the data into test and training portions. They built their model on training data using several different model settings, and report the best test error they achieved.
(b) A project team split their data into training and test. Using their training data and cross-validation, they chose the best parameter setting. They built a model using these parameters and their training data, and then report their error on test data.
登入後即可作答並保存紀錄。
這題考驗對機器學習實驗流程中資料劃分、模型選擇與評估的正確理解,特別是關於資料洩漏(data leakage)和過度擬合的問題。
(a) Problematic。
-
問題一:特徵選擇(Feature Selection)的資料洩漏
- 說明:該團隊在「完整數據集」(full data)上進行了特徵選擇。這意味著他們在劃分訓練集和測試集之前,就已經利用了測試集中的信息來選擇特徵。如果特徵選擇的標準是基於數據的整體表現(例如,與目標變數的相關性,或在某種評估指標上的表現),那麼測試集中的信息就已經「洩漏」到了特徵選擇的過程中。
- 後果:這會導致選擇出的特徵集看似在測試集上表現良好,但實際上是對測試集過擬合。當使用這些「優化」過的特徵集來訓練模型並在獨立的測試集上評估時,性能會比預期差,因為模型實際上已經間接「看到」了測試集的某些特性。
- 正確做法:特徵選擇應該只在訓練集上進行。應先將數據劃分為訓練集和測試集。然後,在訓練集上進行特徵選擇,並使用交叉驗證(cross-validation)在訓練集內部來評估不同特徵子集的效果。最後,在選定最佳特徵子集後,再用這個子集在整個訓練集上訓練模型,並在測試集上進行最終評估。
-
問題二:報告「最佳測試錯誤」(Best Test Error)
- 說明:即使特徵選擇沒有問題,報告「最佳測試錯誤」也暗示著他們可能嘗試了多種模型設置(例如,不同的超參數、不同的模型架構),並選擇了在測試集上表現最好的那個。
- 後果:這類似於在測試集上進行了模型選擇或超參數調優。