標籤:style blog http color io os ar strong for
要解決的問題是,給出了具有2個特徵的一堆訓練資料集,從該資料的分布可以看出它們並不是非常線性可分的,因此很有必要用更高階的特徵來類比。例如本程式中個就用到了特徵值的6次方來求解。
Data
To begin, load the files ‘ex5Logx.dat‘ and ex5Logy.dat‘ into your program. This dataset represents the training set of a logistic regression problem with two features. To avoid confusion later, we will refer to the two input features contained in ‘ex5Logx.dat‘ as and . So in the ‘ex5Logx.dat‘ file, the first column of numbers represents the feature , which you will plot on the horizontal axis, and the second feature represents , which you will plot on the vertical axis.
After loading the data, plot the points using different markers to distinguish between the two classifications. The commands in Matlab/Octave will be:
x = load(‘ex5Logx.dat‘); y = load(‘ex5Logy.dat‘);figure% Find the indices for the 2 classespos = find(y); neg = find(y == 0);plot(x(pos, 1), x(pos, 2), ‘+‘)hold onplot(x(neg, 1), x(neg, 2), ‘o‘)
After plotting your image, it should look something like this:
Model
the hypothesis function is
Let‘s look at the parameter in the sigmoid function .
In this exercise, we will assign to be all monomials (meaning polynomial terms) of and up to the sixth power:
To clarify this notation: we have made a 28-feature vector where
此時加入了規則項後的系統的損失函數為:
Newton’s method
Recall that the Newton‘s Method update rule is
1. is your feature vector, which is a 28x1 vector in this exercise.
2. is a 28x1 vector.
3. and are 28x28 matrices.
4. and are scalars.
5. The matrix following in the Hessian formula is a 28x28 diagonal matrix with a zero in the upper left and ones on every other diagonal entry.
After convergence, use your values of theta to find the decision boundary in the classification problem. The decision boundary is defined as the line where
Code
%載入資料clc,clear,close all;x = load(‘ex5Logx.dat‘);y = load(‘ex5Logy.dat‘);%畫出資料的分布圖plot(x(find(y),1),x(find(y),2),‘o‘,‘MarkerFaceColor‘,‘b‘)hold on;plot(x(find(y==0),1),x(find(y==0),2),‘r+‘)legend(‘y=1‘,‘y=0‘)% Add polynomial features to x by % calling the feature mapping function% provided in separate m-filex = map_feature(x(:,1), x(:,2)); %投影到高維特徵空間[m, n] = size(x);% Initialize fitting parameterstheta = zeros(n, 1);% Define the sigmoid functiong = inline(‘1.0 ./ (1.0 + exp(-z))‘); % setup for Newton‘s methodMAX_ITR = 15;J = zeros(MAX_ITR, 1);% Lambda is the regularization parameterlambda = 1;%lambda=0,1,10,修改這個地方,運行3次可以得到3種結果。% Newton‘s Methodfor i = 1:MAX_ITR % Calculate the hypothesis function z = x * theta; h = g(z); % Calculate J (for testing convergence) -- 損失函數 J(i) =(1/m)*sum(-y.*log(h) - (1-y).*log(1-h))+ ... (lambda/(2*m))*norm(theta([2:end]))^2; % Calculate gradient and hessian. G = (lambda/m).*theta; G(1) = 0; % extra term for gradient L = (lambda/m).*eye(n); L(1) = 0;% extra term for Hessian grad = ((1/m).*x‘ * (h-y)) + G; H = ((1/m).*x‘ * diag(h) * diag(1-h) * x) + L; % Here is the actual update theta = theta - H\grad; end% Plot the results % We will evaluate theta*x over a % grid of features and plot the contour % where theta*x equals zero% Here is the grid rangeu = linspace(-1, 1.5, 200);v = linspace(-1, 1.5, 200);z = zeros(length(u), length(v));% Evaluate z = theta*x over the gridfor i = 1:length(u) for j = 1:length(v) z(i,j) = map_feature(u(i), v(j))*theta;%這裡繪製的並不是損失函數與迭代次數之間的曲線,而是線性變換後的值 endendz = z‘; % important to transpose z before calling contour% Plot z = 0% Notice you need to specify the range [0, 0]contour(u, v, z, [0, 0], ‘LineWidth‘, 2)%在z上畫出為0值時的介面,因為為0時剛好機率為0.5,符合要求legend(‘y = 1‘, ‘y = 0‘, ‘Decision boundary‘)title(sprintf(‘\\lambda = %g‘, lambda), ‘FontSize‘, 14)hold off% Uncomment to plot J% figure% plot(0:MAX_ITR-1, J, ‘o--‘, ‘MarkerFaceColor‘, ‘r‘, ‘MarkerSize‘, 8)% xlabel(‘Iteration‘); ylabel(‘J‘)
Result
Regularized logistic regression