損失函數(Loss Function) -1

來源:互聯網
上載者:User

標籤:des   style   blog   http   io   color   ar   os   使用   

http://www.ics.uci.edu/~dramanan/teaching/ics273a_winter08/lectures/lecture14.pdf

  1. Loss Function

    損失函數可以看做 誤差部分(loss term) + 正則化部分(regularization term)

1.1 Loss Term

  • Gold Standard (ideal case)
  • Hinge (SVM, soft margin)
  • Log (logistic regression, cross entropy error)
  • Squared loss (linear regression)
  • Exponential loss (Boosting)

??

Gold Standard 又被稱為0-1 loss, 記錄分類錯誤的次數

Hinge Losshttp://en.wikipedia.org/wiki/Hinge_loss

For an intended output?t?= ±1?and a classifier score?y, the hinge loss of the prediction?y?is defined as

Note that?y?should be the "raw" output of the classifier‘s decision function, not the predicted class label. E.g., in linear SVMs,?

It can be seen that when?t?and?y?have the same sign (meaning?y?predicts the right class) and?

, the hinge loss?

, but when they have opposite sign,?

increases linearly with?y?(one-sided error).

??

來自 <http://en.wikipedia.org/wiki/Hinge_loss>

Plot of hinge loss (blue) vs. zero-one loss (misclassification, green:y?< 0) for?t?= 1?and variable?y. Note that the hinge loss penalizes predictions?y?< 1, corresponding to the notion of a margin in a support vector machine.

??

來自 <http://en.wikipedia.org/wiki/Hinge_loss>

??

??

在Pegasos: Primal Estimated sub-GrAdient SOlver for SVM論文中

這裡把第一部分看成正規化部分,第二部分看成誤差部分,注意對比ng關於svm的課件

不考慮規則化

考慮規則化

??

Log Loss

Ng的課件1,先是講 linear regression 然後引出最小二乘誤差,之後機率角度高斯分布解釋最小誤差。

然後講羅吉斯迴歸,使用MLE來引出最佳化目標是使得所見到的訓練資料出現機率最大

??

??

最大化下面的log似然函數

而這個恰恰就是最小化cross entropy!

??

http://en.wikipedia.org/wiki/Cross_entropy

http://www.cnblogs.com/rocketfan/p/3350450.html 資訊理論,交叉熵與KL divergence關係

??

Cross entropy can be used to define loss function in machine learning and optimization. The true probability?

?is the true label, and the given distribution?

?is the predicted value of the current model.

More specifically, let us consider?logistic regression, which (in its most basic guise) deals with classifying a given set of data points into two possible classes generically labelled?

?and?

. The logistic regression model thus predicts an output?

, given an input vector?

. The probability is modeled using thelogistic function?

. Namely, the probability of finding the output?

?is given by

where the vector of weights?

?is learned through some appropriate algorithm such as?gradient descent. Similarly, the conjugate probability of finding the output?

?is simply given by

The true (observed) probabilities can be expressed similarly as?

?and?

.

??

Having set up our notation,?

?and?

, we can use cross entropy to get a measure for similarity between?

?and?

:

The typical loss function that one uses in logistic regression is computed by taking the average of all cross-entropies in the sample. For specifically, suppose we have?

?samples with each sample labeled by?

. The loss function is then given by:

where?

, with?

?the logistic function as before.

??

The logistic loss is sometimes called cross-entropy loss. It‘s also known as log loss (In this case, the binary label is often denoted by {-1,+1}).[1]

??

來自 <http://en.wikipedia.org/wiki/Cross_entropy>

??

??

因此和ng從MLE角度給出的結論是完全一致的! 差別是最外面的一個負號

也就是羅吉斯迴歸的最佳化目標函數是 交叉熵

??

squared loss

??

exponential loss

指數誤差通常用在boosting中,指數誤差始終> 0,但是確保越接近正確的結果誤差越小,反之越大。

??

??

損失函數(Loss Function) -1

聯繫我們

該頁面正文內容均來源於網絡整理,並不代表阿里雲官方的觀點,該頁面所提到的產品和服務也與阿里云無關,如果該頁面內容對您造成了困擾,歡迎寫郵件給我們,收到郵件我們將在5個工作日內處理。

如果您發現本社區中有涉嫌抄襲的內容,歡迎發送郵件至: info-contact@alibabacloud.com 進行舉報並提供相關證據,工作人員會在 5 個工作天內聯絡您,一經查實,本站將立刻刪除涉嫌侵權內容。

A Free Trial That Lets You Build Big!

Start building with 50+ products and up to 12 months usage for Elastic Compute Service

  • Sales Support

    1 on 1 presale consultation

  • After-Sales Support

    24/7 Technical Support 6 Free Tickets per Quarter Faster Response

  • Alibaba Cloud offers highly flexible support services tailored to meet your exact needs.