> For the complete documentation index, see [llms.txt](https://json007.gitbook.io/svm/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://json007.gitbook.io/svm/linear_model/guang_yi_xian_xing_mo_xing.md).

# Generalized Linear Models

## 自然指数分布族

自然指数分布族 [Exponential family](https://en.wikipedia.org/wiki/Exponential_family%29%20：%20%20%0A如果一个概率分布可以表示成%20$$p%28y;/eta%29%20=%20b%28y%29%20/exp%28/eta%5ET%20T%28y%29-a%28/eta%29)$$，则称x服从自然指数分布族分布。\
多看看wiki的介绍。涉及到的东西很多。

指数分布族包括：

* [Normal distribution](https://en.wikipedia.org/wiki/Normal_distribution)，多元正态分布
* [Bernoulli distribution](https://en.wikipedia.org/wiki/Bernoulli_distribution%29（01问题建模），\[Categorical%20distribution]%28https://en.wikipedia.org/wiki/Categorical_distribution)（对k个结果的事件建模），
* [Poisson distribution](https://en.wikipedia.org/wiki/Poisson_distribution)（对计数过程建模）
* [Gamma distribution](https://en.wikipedia.org/wiki/Gamma_distribution%29%20，\[Exponential%20distribution]%28https://en.wikipedia.org/wiki/Exponential_distribution)（对实数的间隔问题建模）
* [Beta distribution](https://en.wikipedia.org/wiki/Beta_distribution)（对小数建模）
* [Dirichlet distribution](https://en.wikipedia.org/wiki/Dirichlet_distribution) （对概率分布进行建模）&#x20;
* [Wishart distribution](https://en.wikipedia.org/wiki/Wishart_distribution) （协方差矩阵的分布）&#x20;

### 正太分布

$$
N(x;\mu,\sigma) = \frac {1}{\sqrt {2\pi} \sigma} \exp(-\frac {(x-\mu)^2}{2\sigma^2})
$$

### bernoulli分布

$$
p(y|p) = p^y(1-p)^{1-y} = \exp (y \log \frac {p}{1-p} + \log (1-p)) \\
\text{link function:} \eta(p) = \log \frac {p}{1-p} \\
\text{response function:} p = \frac {1}{1+e^{-\eta}}
$$

### 泊松分布

$$
p(x|\lambda) = \frac {\lambda^\*}{x!} e^{-\lambda} = \exp(x \ln \lambda-\lambda) \frac {1}{x!}
$$

### student分布

$$
T(x|\mu,\sigma,v) = \[1+\frac {1}{v}(\frac {x-\mu} {\sigma})^2]^{- \frac {v+1}{2}}
$$

[学生t-分布](https://zh.wikipedia.org/wiki/学生t-分布)

> 此分布形式上与高斯分布类似，弥补了高斯分布的一个不足，就是高斯分布对离群的数据非常敏感，但是Student t分布更鲁棒。一般设置ν=4，在大多数实际问题中都有很好的性能，当ν大于等于5时将会是去鲁棒性，同时会迅速收敛到高斯分布。\
> 特别的，当ν=1时，被称为柯西分布（Cauchy）。

Gamma 分布 [怎么来理解伽玛（gamma）分布？](https://www.zhihu.com/question/34866983)

**为什么弄出个指数分布族？**&#x4D;LAPP page313

* 指数分布族理论上都有共轭先验分布
* 将分布全部转换成指数形式，然后给定的约束条件下，熵值最大的函数就是他们各自的分布
* It can be shown that, under certain regularity conditions, the exponential family is the only

  family of distributions with finite-sized sufficient statistics, meaning that we can compress

  the data into a fixed-sized summary without loss of information. This is particularly useful

  for online learning, as we will see later. &#x20;
* The exponential family is the only family of distributions for which **conjugate priors** exist,

  which simplifies the computation of the posterior (see Section 9.2.5).
* The exponential family can be shown to be the family of distributions that makes the least

  set of assumptions subject to some user-chosen constraints (see Section 9.2.6).
* The exponential family is at the core of **generalized linear models**, as discussed in Section 9.3. &#x20;
* The exponential family is at the core of **variational inference**, as discussed in Section 21.2.&#x20;
* **指数簇分布的最大熵**等价于其**指数形式的最大似然**。

## 广义线性模型

广义线性模型，是为了克服线性回归模型的缺点出现的，是线性回归模型的推广。\
首先**自变量可以是离散**的，也可以是连续的。离散的可以是0-1变量，也可以是多种取值的变量。\
与线性回归模型相比较，有以下推广：

* 随机误差项不一定服从正态分布，可以服从二项、泊松、负二项、正态、伽马、逆高斯等分布，这些分布被统称为指数分布族。
* 引入联接函数g(⋅)。因变量和自变量通过联接函数产生影响，即Y=g(Xβ)，联接函数满足单调，可导。常用的联接函数有恒等 $$Y=X\beta$$，对数$$Y=\ln(X\beta)$$，幂函数$$Y=(X\beta)^k$$，平方根$$Y=\sqrt {X\beta}$$，$$Y= logit(\ln(\frac {Y}{1-Y})) = X\beta$$等。 &#x20;

根据不同的数据，可以自由选择不同的模型。大家比较熟悉的Logit模型就是使用Logit联接、随机误差项服从二项分布得到模型。

### three assumptions

* p(y|x;θ)满足指数分布族，也就是说，给定x和θ，y的分布情况满足以η为参数的指数分布族的分布。
* 给定x，我们的目标是预测T(y)的期望值，也即hθ(x)=E\[T(y)|x]
* 自然参数η和输入x是线性关系:η=θTx
* y | x; θ ∼ ExponentialFamily(η). I.e., given x and θ, the distribution of\
  y follows some exponential family distribution, with parameter η.
* Given x, our goal is to predict the expected value of T(y) given x.\
  In most of our examples, we will have T(y) = y, so this means we\
  would like the prediction h(x) output by our learned hypothesis h to\
  25\
  satisfy h(x) = E\[y|x]. (Note that this assumption is satisfied in the\
  choices for hθ(x) for both logistic regression and linear regression. For\
  instance, in logistic regression, we had hθ(x) = p(y = 1|x; θ) = 0 · p(y =\
  0|x; θ) + 1 · p(y = 1|x; θ) = E\[y|x; θ].)
* The natural parameter η and the inputs x are related linearly: η = θ\
  T x.\
  (Or, if η is vector-valued, then ηi = θ\
  T\
  i x.)

对于广义线性模型，取决于采用什么分布。采用正太分布，则得到**最小二乘模型**。采用伯努利分布，则得到**logistic模型**。然后用梯度下降等求线性部分参数。

### 建模

![](https://2270971654-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-M7DcNFhVrwIk3Tks_pB%2Fsync%2F3b0b6b7522e7302e143a72b8645727fa479baa19.png?generation=1589383930237130\&alt=media)

### 参考佳文

[广义线性模型](https://mp.weixin.qq.com/s?__biz=MzA4NDEyMzc2Mw==\&mid=2649677334\&idx=3\&sn=9fccb5c53c4be9039425e93c1a7e122e)

[广义线性模](https://www.zhihu.com/question/28469421%29%20%20%0A\[GLM,%20NON-LINEARITY%20AND%20HETEROSCEDASTICITY]%28http://freakonometrics.hypotheses.org/9593%29%20%20%0A\[一般线性模型、混合线性模型、广义线性模型]%28http://bbs.pinggu.org/thread-2996069-1-1.html%29%20%20%0A\[从线性模型到广义线性模型（1）——模型假设篇]%28http://cos.name/2011/01/how-does-glm-generalize-lm-assumption/%29%20%20%0A\[从线性模型到广义线性模型%282%29——参数估计、假设检验]%28http://cos.name/2011/01/how-does-glm-generalize-lm-fit-and-test/)

[统一分布：指数模型家族](https://zhuanlan.zhihu.com/p/148776108)
