Downhill 0/0
Computational methods of optimization, from zero

How computers find the best answer, one step downhill at a time

Fitting a model to data, routing a delivery van, choosing a diet, training a neural network: each one comes down to finding the lowest point of a function. This course teaches you how that search works and why it works, starting from school algebra and going all the way to conjugate gradients and quasi-Newton methods.

Based on E0 230 at IISc, taught by Prof. Chiranjib Bhattacharyya. An independent student resource, not an official course page: see the acknowledgement.

Starts from school algebra Intuition first, then the formal maths Every symbol defined before it's used Hands-on plots you can drag Proofs are there when you want them
Read, then doEach chapter explains an idea in plain words, shows a picture, and then lets you change something and watch what happens.
You grade yourselfNobody else sees your answers. The practice problems and quizzes are real exam-level questions, so if you can answer them, you understand the topic.
Stuck? That's normalA wrong answer gives you a hint. A second wrong answer, or one click, gives the full worked solution.
"Go deeper" is optionalProofs and extra theory sit in purple dashed boxes. Everything you need for the exam stays in the main path.

Preparing for an exam? Part 12 has practice problems in the professor's style, adapted from past papers.

What's inside: fourteen parts of 1–2.5 hours each, split into short chapters. Go at your own pace; your progress is saved in this browser.

Before any algorithm, we need a shared language. This part teaches you to read the notation used everywhere in optimization, to think of a list of numbers as an arrow in space, and to picture a function of two variables as a landscape. If you already know these, skim the quizzes: if you get them right first time, move on.

You need: school algebra (solving $2x+3=7$, expanding $(a+b)^2$) and a feel for what a graph of $y=x^2$ looks like. Nothing else.

Reading maths notation without fear

Mathematical notation is shorthand for sentences you already understand, and every optimization problem is written in the same few symbols.

Textbooks and exams in this course write everything in this shorthand. Once you can read it aloud, half the difficulty of a theorem disappears.

Learning the abbreviations in a recipe: "tbsp", "°C" and "simmer 5 min" look cryptic until someone tells you what they mean, then you read them without thinking.

Here is a sentence you might see in the first lecture:

$$\min_{\x\in\R^n}\ f(\x)\quad\text{subject to}\quad g_i(\x)\le 0,\ i=1,\dots,m.$$

Read aloud, it says: "Find the smallest value of the function $f$, over all lists $\x$ of $n$ real numbers, among those lists that make every $g_i(\x)$ at most zero." That's all. The rest of this chapter teaches you each piece, so you can do this translation yourself.

Sets: collections of things

A set is a collection of objects, called its elements. We write sets with curly braces: $\{1, 2, 3\}$ is the set containing 1, 2 and 3. Order and repetition don't matter: $\{3,1,2,1\}$ is the same set.

Often we describe a set by a rule instead of a list. The notation $\{x : x^2 < 4\}$ reads "the set of all $x$ such that $x^2<4$". The colon (sometimes a bar, $\{x \mid x^2<4\}$) means "such that".

Symbols for sets and numbers
$x\in S$
"$x$ is an element of $S$" (or "$x$ belongs to $S$"). $x\notin S$ means it isn't.
$A\subseteq B$
"$A$ is a subset of $B$": every element of $A$ is also in $B$.
$A\cap B$, $A\cup B$
Intersection (elements in both) and union (elements in either).
$\emptyset$
The empty set, which has no elements.
$\N,\ \Z,\ \Q$
Natural numbers $\{1,2,3,\dots\}$, integers $\{\dots,-1,0,1,\dots\}$, rationals (fractions $p/q$ with integers $p,q$ and $q\ne0$).
$\R$
The real numbers: every point on the number line, including $\sqrt2$ and $\pi$.
$[a,b]$, $(a,b)$
Closed interval $\{x: a\le x\le b\}$ (endpoints included) and open interval $\{x : a\lt x\lt b\}$ (endpoints excluded). $[a,b)$ includes $a$ but not $b$.
$\R^n$
All lists $(x_1,\dots,x_n)$ of $n$ real numbers. $\R^2$ is the plane, $\R^3$ is space.

Functions: machines that take an input and give one output

A function $f: A\to B$ takes each element of the set $A$ (the domain) and gives back exactly one element of $B$. In this course you'll mostly see $f:\R^n\to\R$: the input is a list of $n$ numbers and the output is a single number. That single number is the thing we want to make small: a cost, an error, an energy.

Example: $f(x_1,x_2) = x_1^2 + 3x_2^2$ takes the pair $(1,2)$ to $1+12=13$.

Sums, quantifiers and arrows

Shorthand you'll meet in every proof
$\sum_{i=1}^{n} a_i$
"The sum of $a_i$ as $i$ goes from 1 to $n$" $= a_1+a_2+\dots+a_n$.
$\forall$
"for all" (or "for every"). $\forall x\in\R,\ x^2\ge 0$.
$\exists$
"there exists" (at least one). $\exists x\in\R : x^2 = 2$.
$P\Rightarrow Q$
"if $P$ then $Q$" ($P$ implies $Q$). Not the same as $Q\Rightarrow P$!
$P\iff Q$
"$P$ if and only if $Q$": each implies the other, so they are equivalent.
$:=$
"is defined to be". $g(x):=x^2$ introduces a new name; it isn't a claim to check.
$\varepsilon,\ \delta$
Greek letters used for "a small positive number". $\alpha,\beta,\lambda,\mu$ are also just names.
A habit that helps: whenever you meet a formula, read it out loud in words. If you can't, stop and look up the symbol you're stuck on. Mathematicians do this too; they're just faster at it.

min, argmin, and the minimizer $\x^\star$

Two different questions hide in "minimize $f$":

  • What is the smallest value? That's $\min_{\x} f(\x)$, a number, often called $f^\star$.
  • Where is it reached? That's $\argmin_{\x} f(\x)$, the set of inputs where $f$ takes that smallest value. A point in this set is called a minimizer and written $\x^\star$ ("x star").

For $f(x)=(x-3)^2+1$: the minimum value is $1$, and $\argmin f=\{3\}$, so $x^\star=3$ and $f^\star=f(x^\star)=1$.

Sometimes there is no smallest value at all, only a value the function gets closer and closer to. For $f(x)=e^{x}$ on $\R$, the outputs approach 0 but never reach it. We then say the infimum is $0$ (written $\inf f = 0$) but the minimum doesn't exist. Part 2 makes this precise; for now just notice that "min" is a promise that the lowest value is actually reached.

Given an objective $f:\R^n\to\R$ and constraint functions $g_i$, $h_j$, the problem $$\min_{\x\in\R^n} f(\x)\quad\text{s.t.}\quad g_i(\x)\le 0\ (i=1,\dots,m),\qquad h_j(\x)=0\ (j=1,\dots,p)$$ asks for a point in the feasible set $\mathcal F=\{\x : g_i(\x)\le0\ \forall i,\ h_j(\x)=0\ \forall j\}$ with the smallest value of $f$. If there are no constraints ($\mathcal F=\R^n$), the problem is unconstrained. "s.t." means "subject to".

Try it

Slide the limits of the interval $[a,b]$ and watch where $f(x) = (x-1)^2$ is smallest on it. Notice when the minimizer sits inside the interval and when it gets pushed to an edge.

Translate into words, and solve: $\displaystyle \min_{x\in[0,2]} (x-3)^2$.

  1. In words: "Find the smallest value of $(x-3)^2$ over all real $x$ between 0 and 2, endpoints included."

    The subscript under "min" tells you where you're allowed to search: here, the closed interval $[0,2]$.

  2. Without constraints, the minimizer would be $x=3$, where the square is 0. But $3\notin[0,2]$.

    A square is never negative and equals 0 only when the inside is 0, so the unconstrained answer is easy. Always check whether it's allowed.

  3. On $[0,2]$ the function $(x-3)^2$ gets smaller as $x$ moves towards 3, so it is smallest at the right edge, $x^\star=2$.

    On this interval, $x-3$ is negative and its size $3-x$ shrinks as $x$ grows, so the square shrinks too.

  4. Answer: $\min = (2-3)^2 = 1$, and $\argmin=\{2\}$.

    Always report both the value and where it's attained if asked; they're different objects.

Compute $\displaystyle\sum_{i=1}^{4} (-1)^{i+1}\, i^2$.

Write out the four terms: $i=1,2,3,4$. The factor $(-1)^{i+1}$ is $+1$ when $i$ is odd and $-1$ when $i$ is even.

$1^2 - 2^2 + 3^2 - 4^2 = 1-4+9-16=-10$.

Find $\min_{x\in\R} f(x)$ and $\argmin_{x\in\R} f(x)$ for $f(x) = x^2 - 2x - 1$.

Complete the square: $x^2-2x = (x-1)^2 - 1$.

$f(x)=(x-1)^2-2$. The square is $\ge 0$ and equals 0 only at $x=1$. So the minimum value is $-2$, reached only at $x^\star=1$, i.e. $\argmin f = \{1\}$.

Let $f(x)=1/x$ on the set $x\in[1,\infty)$. The function has no minimum. What is its infimum $\inf_{x\ge1} f(x)$?

What happens to $1/x$ as $x$ becomes huge? Can it ever become negative?

$1/x>0$ for every $x\ge1$, so $0$ is a lower bound. And $1/x$ gets as close to 0 as you like (take $x=10^6$), so no bigger number is a lower bound. Hence $\inf f=0$, but no $x$ has $f(x)=0$, so the minimum doesn't exist.

  • Read every formula aloud in words before trying to use it
  • Check where the "min" is allowed to search (the set under it)
  • Distinguish the value $f^\star$ from the point $\x^\star$
  • Treat $:=$ as a definition, not a fact to prove
  • Confusing $P\Rightarrow Q$ with $Q\Rightarrow P$
  • Writing "min" when the lowest value is never actually reached (use "inf")
  • Forgetting whether an interval's endpoints are included: $[a,b]$ vs $(a,b)$
  • Answering "where" when asked "what value", or vice versa
  1. Optimization problems are written as "min $f(\x)$ over a set": find the input with the smallest output.
  2. $\min$ is a value; $\argmin$ is the set of points achieving it; $\x^\star$ is one such point.
  3. Some functions have no minimum, only an infimum that is approached but never reached.

What does $\{x\in\R : x^2\le 9\}$ describe?

The numbers 9, 4, 1, 0
That lists some squares, but the set is about which $x$ have $x^2\le 9$. Read the colon as "such that".
The closed interval $[-3,3]$
$x^2\le 9$ exactly when $-3\le x\le 3$, endpoints included because $(\pm3)^2=9$ satisfies "$\le$".
The open interval $(-3,3)$
Close. Check the endpoints: does $x=3$ satisfy $x^2\le 9$?

For $f(x)=(x+2)^2+5$, which statement is correct?

$\min f = -2$
$-2$ is where the minimum happens, not the smallest value. What is $f(-2)$?
$\argmin f = 5$
5 is a value of $f$; argmin is a set of inputs.
$\min f = 5$ and $\argmin f=\{-2\}$
The square is smallest (zero) at $x=-2$, where $f=5$.

"$\forall \varepsilon>0\ \exists N$ such that …" is read as:

For every positive $\varepsilon$ there is some $N$ such that …
$\forall$ is "for every", $\exists$ is "there exists". The order matters: $N$ is allowed to depend on $\varepsilon$.
There is an $N$ that works for every positive $\varepsilon$
That swaps the order of the quantifiers, which is a much stronger claim. Read left to right.
For some $\varepsilon$ there exists $N$…
$\forall$ means "for all", not "for some".

Which problem is unconstrained?

$\min x_1+x_2$ s.t. $x_1\ge 0$
"s.t. $x_1\ge0$" is a constraint: it restricts where we may search.
$\min_{\x\in\R^2} (x_1-1)^2 + x_2^4$
We may search over all of $\R^2$, so there are no constraints.
$\min_{x\in[0,1]} x^2$
The interval $[0,1]$ is a constraint: it's the same as $0\le x$ and $x\le 1$.

Vectors: lists of numbers that are also arrows

A vector is a list of numbers that you can also picture as an arrow, and the dot product measures how much two arrows point the same way.

Every algorithm in this course moves a point in the direction of some vector. Whether a step goes "downhill" is decided by the sign of a dot product.

Directions on a map: "3 blocks east, 2 blocks north" is a list of two numbers, and also an arrow from where you stand to where you end up.

A vector in $\R^n$ is a list of $n$ real numbers, written as a column: $$\x = \begin{pmatrix}x_1\\ x_2\\ \vdots\\ x_n\end{pmatrix}\in\R^n.$$ To save space we often write it as a row with a "transpose" mark, $\x=(x_1,\dots,x_n)^\top$. We print vectors in bold ($\x$, $\d$, $\g$) and plain numbers (also called scalars) in normal type ($\alpha$, $t$, $x_1$).

In $\R^2$ you can draw $\x=(3,2)^\top$ as an arrow from the origin to the point $(3,2)$. The same list can mean a position (a point) or a direction and distance (an arrow); context tells you which.

Adding and scaling

Add vectors entry by entry: $(1,2)^\top+(3,-1)^\top=(4,1)^\top$. As arrows: put the second arrow's tail at the first arrow's head. Multiply by a scalar to stretch: $2(1,2)^\top=(2,4)^\top$, and $-1$ flips the direction.

This is what every algorithm in this course does: $$\x_{k+1} = \x_k + \alpha_k \d_k,$$ "from the current point $\x_k$, move a distance controlled by $\alpha_k$ in the direction $\d_k$". The subscript $k$ counts the steps: $\x_0$ is where we start, $\x_1$ is after one step, and so on.

For $\x,\y\in\R^n$, the inner product (or dot product) is the number $$\ip{\x}{\y} = \x^\top\y = \sum_{i=1}^n x_i y_i = x_1y_1+\dots+x_ny_n.$$ The Euclidean norm (length) of $\x$ is $\norm{\x}=\sqrt{\ip{\x}{\x}}=\sqrt{x_1^2+\dots+x_n^2}$.

Example: $\x=(3,4)^\top$, $\y=(1,2)^\top$. Then $\x^\top\y=3+8=11$ and $\norm{\x}=\sqrt{9+16}=5$, which is Pythagoras' theorem.

What the dot product means

The dot product has a geometric meaning that makes it the most important operation in this course: $$\x^\top\y = \norm{\x}\,\norm{\y}\cos\theta,$$ where $\theta$ is the angle between the arrows. So:

  • $\x^\top\y>0$: the arrows point "roughly the same way" (angle less than 90°).
  • $\x^\top\y=0$: they are perpendicular (also called orthogonal).
  • $\x^\top\y<0$: they point "roughly opposite ways" (angle more than 90°).

Later you'll learn that a direction $\d$ goes downhill from $\x$ whenever $\d^\top\grad f(\x)<0$, where $\grad f$ is the gradient. Everything starts here.

Try it

Drag the tips of the two arrows. Make the dot product exactly zero (perpendicular), then negative. Watch the shadow (projection) of $\y$ onto $\x$ flip to the other side.

For all $\x,\y\in\R^n$: $\quad|\x^\top\y|\le\norm{\x}\,\norm{\y}$, with equality exactly when one vector is a multiple of the other.

In words: the dot product can never exceed the product of the lengths. That's just $|\cos\theta|\le1$ in disguise. It has a consequence we'll use for gradient descent: among all directions $\d$ of length 1, the one that makes $\d^\top\g$ as negative as possible is $\d=-\g/\norm{\g}$, pointing exactly opposite to $\g$.

A second consequence, the triangle inequality $\norm{\x+\y}\le\norm{\x}+\norm{\y}$, says that going straight is never longer than taking a detour.

Go deeper: a two-line proof of Cauchy–Schwarz

If $\y=\0$ both sides are 0. Otherwise, for every real $t$, a squared length is never negative: $$0\le\norm{\x-t\y}^2=\norm{\x}^2-2t\,\x^\top\y+t^2\norm{\y}^2.$$ This is a quadratic in $t$ that never goes below zero, so it has at most one real root, which means its discriminant is $\le0$: $4(\x^\top\y)^2-4\norm{\x}^2\norm{\y}^2\le0$. Rearranging gives $|\x^\top\y|\le\norm{\x}\norm{\y}$. Equality means the quadratic touches zero, i.e. $\x=t\y$ for some $t$.

Other ways to measure length

The Euclidean norm is not the only sensible "size". Two others appear in optimization:

  • $\norm{\x}_1=|x_1|+\dots+|x_n|$, the "taxicab" length (blocks walked on a grid).
  • $\norm{\x}_\infty=\max_i |x_i|$, the largest single entry.

Any function that is positive for nonzero vectors, scales as $\norm{c\x}=|c|\norm{\x}$, and obeys the triangle inequality is called a norm. The set of vectors with norm at most 1 is the norm's unit ball: it looks different for each norm.

Try it

Switch between the norms. The shape is all points with "length" at most 1. Drag the point to see its three lengths side by side.

Let $\g=(2,-1)^\top$. Which of $\d_1=(1,1)^\top$, $\d_2=(-1,0)^\top$, $\d_3=(1,2)^\top$ make a negative dot product with $\g$? Which unit vector makes it as negative as possible?

  1. $\g^\top\d_1 = 2-1 = 1>0$, $\ \g^\top\d_2=-2<0$, $\ \g^\top\d_3=2-2=0$.

    Multiply matching entries and add. The sign tells you whether the angle with $\g$ is below, at, or above 90°.

  2. Only $\d_2$ has a negative dot product. $\d_3$ is perpendicular to $\g$.

    Zero means exactly 90°: such a direction is "sideways" relative to $\g$.

  3. By Cauchy–Schwarz, for a unit vector $\u$, $\g^\top\u\ge-\norm{\g}=-\sqrt5$, with equality when $\u$ points opposite to $\g$: $\u=-\g/\norm{\g}=(-2,1)^\top/\sqrt5$.

    Equality in Cauchy–Schwarz needs the vectors to be parallel; to make the product negative, they must point opposite ways.

Compute $\x^\top\y$ for $\x=(1,-2,3)^\top$ and $\y=(4,0,-3)^\top$.

Multiply entry by entry: $1\cdot4$, $(-2)\cdot0$, $3\cdot(-3)$, then add.

$4+0-9=-5$.

For $\x=(5,-12)^\top$, compute $\norm{\x}_2$, $\norm{\x}_1$ and $\norm{\x}_\infty$. Careful with the second one.

$\norm{\x}_2=\sqrt{25+144}$. For $\norm{\x}_1$, add absolute values: $|5|+|-12|$.

$\norm{\x}_2=\sqrt{169}=13$; $\norm{\x}_1=5+12=17$; $\norm{\x}_\infty=\max(5,12)=12$. If you got 7 for the 1-norm, you added $5+(-12)$ without taking absolute values.

For which value of $t$ are $\u=(t,1)^\top$ and $\v=(1,-2)^\top$ perpendicular?

Perpendicular means $\u^\top\v=0$. Write $\u^\top\v$ in terms of $t$.

$\u^\top\v=t-2=0$, so $t=2$.

What is the angle (in degrees) between $(1,1,0)^\top$ and $(1,-1,5)^\top$?

Compute the dot product first. You may not need the lengths at all.

The dot product is $1-1+0=0$, so $\cos\theta=0$ and $\theta=90^\circ$.

  • Picture vectors as arrows when reasoning about directions
  • Use the sign of $\x^\top\y$ to judge "same way / sideways / opposite"
  • Remember $\norm{\x}^2=\x^\top\x$: it removes square roots from algebra
  • Forgetting absolute values in $\norm{\cdot}_1$
  • Thinking $\x^\top\y=0$ means one of them is zero (it usually means perpendicular)
  • Writing $\norm{\x+\y}=\norm{\x}+\norm{\y}$; it's only $\le$ in general
  1. A vector is a list of numbers and an arrow; algorithms move points by adding scaled arrows: $\x_{k+1}=\x_k+\alpha_k\d_k$.
  2. $\x^\top\y=\norm{\x}\norm{\y}\cos\theta$: its sign tells you whether two directions agree.
  3. Cauchy–Schwarz: $|\x^\top\y|\le\norm{\x}\norm{\y}$, and the most negative $\g^\top\u$ over unit $\u$ is at $\u=-\g/\norm{\g}$.

If $\a^\top\b<0$, the angle between $\a$ and $\b$ is…

less than 90°
That's when the dot product is positive, since $\cos\theta>0$ there.
exactly 90°
At exactly 90°, $\cos\theta=0$ and so is the dot product.
more than 90°
$\a^\top\b=\norm{\a}\norm{\b}\cos\theta<0$ forces $\cos\theta<0$, i.e. an obtuse angle.

Which equals $\norm{\x}^2$?

$\norm{\x}_1^2$
The 1-norm adds absolute values; squaring it gives cross terms like $2|x_1||x_2|$.
$\x^\top\x$
$\x^\top\x=\sum x_i^2$, the square of the Euclidean length.
$\sum_i x_i$
That can even be negative: try $\x=(-1,0)$.

By Cauchy–Schwarz, if $\norm{\g}=4$ and $\norm{\u}=1$, the smallest possible value of $\g^\top\u$ is:

0
We can do better: point $\u$ against $\g$.
−1
The lengths matter: $|\g^\top\u|$ can be as large as $\norm{\g}\norm{\u}$.
−4
$\g^\top\u\ge-\norm{\g}\norm{\u}=-4$, reached at $\u=-\g/4$.

The unit ball of the $\infty$-norm in $\R^2$ is a…

circle
That's the Euclidean (2-norm) ball.
diamond (a square rotated 45°)
That's the 1-norm ball, $|x_1|+|x_2|\le1$.
square with sides parallel to the axes
$\max(|x_1|,|x_2|)\le1$ means both $|x_1|\le1$ and $|x_2|\le1$: the square $[-1,1]^2$.

Functions of two variables as landscapes

A function of two variables is a landscape: the inputs are your position on a map, the output is the height, and contour lines join points of equal height.

Every interactive plot in this course is a contour map. Reading them fluently lets you see why an algorithm zig-zags or races to the answer.

A hiking map: you never see the mountain itself, only the contour lines, yet you can tell where it's steep (lines close together) and where the valley floor is.

Our running example: fit the line

You measured three data points: $(1,2)$, $(2,3)$ and $(3,5)$. You want the straight line $y = wx+c$ that fits them best. Each choice of slope $w$ and intercept $c$ makes some errors. Square each error and add them up: $$L(w,c)=(w\cdot1+c-2)^2+(w\cdot2+c-3)^2+(w\cdot3+c-5)^2.$$ $L$ is a function of two variables, $(w,c)$. The best line is the minimizer of $L$. This little problem, "least squares", will follow us through the whole course: gradient descent, conjugate gradients and Newton's method will all be tested on it.

Try it

Drag the dot on the right-hand map (or use the sliders) to choose a slope $w$ and intercept $c$. The left panel shows the line and its errors as squares; the right shows the landscape $L(w,c)$ as contour lines. Try to find the bottom of the valley.

Graphs and level sets

For $f:\R^2\to\R$, the graph is the surface of points $(x_1,x_2,f(x_1,x_2))$ in 3-D. Drawing surfaces is hard, so we use the map view instead.

For a number $c$, the level set of $f$ at height $c$ is $\{\x : f(\x)=c\}$, all inputs where $f$ equals $c$. A sublevel set is $\{\x : f(\x)\le c\}$, everything at height $c$ or below.

For $f(\x)=x_1^2+x_2^2$, the level set at height $c>0$ is a circle of radius $\sqrt c$. For $f(\x)=x_1^2+4x_2^2$, the level sets are ellipses that are twice as wide as they are tall. How stretched those ellipses are turns out to control how fast gradient descent runs: that's the condition number you'll meet in Part 0b and Part 5.

Three ways to read a contour map:

  • Lines close together mean the height changes fast there: steep ground.
  • Closed loops shrinking to a point mark a pit (a minimum) or a peak.
  • A contour that crosses itself in an X marks a saddle, like a mountain pass. You'll study these in Part 3.
Try it

Pick a function and drag the height slider. The highlighted curve is the level set at that height; the shaded region is the sublevel set. Watch how the shape changes for the stretched bowl and the saddle.

Describe the level sets of $f(x_1,x_2)=x_1^2+4x_2^2$ at heights $c=4$ and $c=16$.

  1. Set $x_1^2+4x_2^2=4$ and divide by 4: $\dfrac{x_1^2}{4}+\dfrac{x_2^2}{1}=1$.

    Writing the equation as "something $=1$" puts it in the standard form of an ellipse $x_1^2/a^2+x_2^2/b^2=1$.

  2. That's an ellipse with half-width $a=2$ along $x_1$ and half-height $b=1$ along $x_2$.

    Setting $x_2=0$ gives $x_1=\pm2$; setting $x_1=0$ gives $x_2=\pm1$.

  3. For $c=16$: $\dfrac{x_1^2}{16}+\dfrac{x_2^2}{4}=1$, the same shape scaled by 2 (half-axes 4 and 2).

    For a quadratic, quadrupling the height doubles every distance, because $f(2\x)=4f(\x)$.

  4. All level sets are ellipses with axis ratio $2:1$; they are nested, shrinking to the minimizer $\0$.

    The ratio of axis lengths is $\sqrt{4/1}=2$: the square root of the ratio of the coefficients. Remember this when we meet eigenvalues.

The level set of $f(x_1,x_2)=x_1^2+x_2^2$ at height $c=9$ is a circle. What is its radius?

$x_1^2+x_2^2=r^2$ is a circle of radius $r$.

$r^2=9$ so $r=3$.

In the fit-the-line example, compute $L(w,c)$ at $w=1.5$, $c=1/3$. (This turns out to be the best line.)

The predictions are $1.5+1/3$, $3+1/3$, $4.5+1/3$. Subtract the data values $2,3,5$, square, and add.

Errors: $\tfrac{11}{6}-2=-\tfrac16$, $\ \tfrac{10}{3}-3=\tfrac13$, $\ \tfrac{29}{6}-5=-\tfrac16$. Squares: $\tfrac1{36}+\tfrac4{36}+\tfrac1{36}=\tfrac{6}{36}=\tfrac16\approx0.1667$.

What shape are the level sets $\{\x: 9x_1^2+x_2^2=c\}$ for $c>0$?

Divide by $c$ and compare with $x_1^2/a^2+x_2^2/b^2=1$.

$x_1^2/(c/9)+x_2^2/c=1$: ellipses with half-axes $\sqrt{c}/3$ and $\sqrt c$, three times taller than wide.

  • Read contour maps for steepness (spacing) and for minima (shrinking loops)
  • Turn "find the best fit" into "minimize a sum of squared errors"
  • Check level sets of quadratics by setting one variable to 0
  • Confusing the level set $f=c$ (a curve) with the sublevel set $f\le c$ (a region)
  • Assuming equally spaced contour lines mean equally spaced heights in every plot (check the levels)
  • Thinking a 2-variable function needs a 3-D picture; a contour map is usually clearer
  1. A function $f:\R^2\to\R$ is a landscape; its contour map shows level sets $\{\x:f(\x)=c\}$.
  2. Fitting a line by least squares means minimizing a bowl-shaped function $L(w,c)$ of two variables.
  3. Stretched elliptical contours signal a bowl that is steep one way and flat the other; this will matter a lot for algorithms.

On a contour map, contour lines are tightly packed in one region. This means…

the function changes quickly there (steep)
Each line is a fixed step in height; packing many into a short distance means a big change over a small move.
the function is nearly constant there
That's the opposite: flat regions have widely spaced lines.
there must be a minimum there
Minima appear as small closed loops at the centre of nested contours, not as dense spacing.

The sublevel set $\{\x: x_1^2+x_2^2\le 4\}$ is…

the circle of radius 2
The circle is the level set ($=4$). "$\le$" includes the inside too.
the filled disc of radius 2
All points within distance 2 of the origin, including the boundary circle.
the disc of radius 4
The radius is $\sqrt4=2$, not 4.

Why is the fit-the-line loss $L(w,c)$ a function of two variables, even though the data have six numbers?

Because there are two axes, $x$ and $y$
The axes describe the data. The loss depends on what we're free to choose.
Because we choose two numbers, $w$ and $c$; the data are fixed constants
The data are given. The unknowns we optimize over, the decision variables, are the slope and intercept.
Because there are two errors
There are three errors, one per data point.

Which pair is correct for $f(x)=4-(x-1)^2$ on $\R$?

$\min f=4$, $\argmin f=\{1\}$
Look at the sign: $-(x-1)^2$ is at most 0, so 4 is the largest value.
The minimum does not exist; $\inf f = -\infty$
As $|x|$ grows, $-(x-1)^2$ goes to $-\infty$, so there's no lowest value at all.
$\min f=0$
$f$ takes negative values, e.g. $f(5)=4-16=-12$.

Which unit vector $\u$ makes $\g^\top\u$ smallest, for $\g=(0,-3)^\top$?

$(0,-1)^\top$
That points the same way as $\g$: $\g^\top\u=3$, the largest.
$(0,1)^\top$
$-\g/\norm{\g}=(0,3)/3=(0,1)$, giving $\g^\top\u=-3$.
$(1,0)^\top$
Perpendicular to $\g$: the product is 0, not the smallest.

Level sets of $f(\x)=x_1^2+100x_2^2$ are ellipses. How many times wider (along $x_1$) than tall (along $x_2$) are they?

100
The axis ratio is the square root of the coefficient ratio.
10
$x_1^2+100x_2^2=c$ has half-axes $\sqrt c$ and $\sqrt{c}/10$: ratio $\sqrt{100}=10$.
1
Equal coefficients would give circles; here they differ by a factor of 100.

"$\x^\star\in\argmin_{\x\in\mathcal F} f(\x)$" means:

$\x^\star$ is feasible and $f(\x^\star)\le f(\x)$ for every feasible $\x$
That's a global minimizer over $\mathcal F$.
$f(\x^\star)$ is the smallest value of $f$ on all of $\R^n$
The search is restricted to $\mathcal F$; outside it, $f$ could be even smaller.
$\x^\star$ is the only point with the smallest value
argmin is a set; it may contain several points. "$\in$" says $\x^\star$ is one of them.

Optimization algorithms see a function only through its slopes (first derivatives) and its curvature (second derivatives). This part builds both from the one-variable derivative you know from school, and introduces the matrix tools (eigenvalues and positive definiteness) that measure curvature in many directions at once.

You need: Part 0a, and the school derivative: the slope of $x^2$ is $2x$, of $x^3$ is $3x^2$, of $e^x$ is $e^x$. The chain rule $\frac{d}{dx}g(h(x))=g'(h(x))h'(x)$ helps but we'll recall it.

The gradient: the slope in every direction at once

The gradient $\grad f(\x)$ is a vector of slopes that tells you how fast $f$ changes in any direction, and it points straight uphill.

Gradient descent, the first algorithm of the course, is literally "step against the gradient". Its justification is a two-line proof you'll be able to write after this chapter.

Standing on a hillside in fog: you can't see the valley, but your feet feel which way the ground tilts. The gradient is that feeling, written as numbers.

Recap: the derivative in one variable

For $f:\R\to\R$, the derivative $f'(x)$ is the slope of the graph at $x$: $$f'(x)=\lim_{h\to0}\frac{f(x+h)-f(x)}{h}.$$ "$\lim_{h\to0}$" means "the value this ratio settles down to as $h$ gets closer and closer to 0". The derivative gives the best straight-line approximation near $x$: $$f(x+h)\approx f(x)+f'(x)\,h\qquad\text{for small }h.$$ If $f'(x)>0$, nudging $x$ to the right increases $f$; if $f'(x)<0$, it decreases $f$. If $f'(x)=0$, the graph is flat there, which is where minima can hide.

Many variables: one slope per coordinate

Now take $f:\R^n\to\R$. Freeze every coordinate except $x_i$, and you have an ordinary one-variable function. Its slope is the partial derivative $\frac{\partial f}{\partial x_i}$ ("partial f, partial x i"). You compute it exactly like a school derivative, treating the other variables as constants.

Example: $f(x_1,x_2)=x_1^2+3x_1x_2$. Then $\frac{\partial f}{\partial x_1}=2x_1+3x_2$ (treat $x_2$ as a number) and $\frac{\partial f}{\partial x_2}=3x_1$ (treat $x_1$ as a number).

The gradient of $f$ at $\x$ is the vector of all partial derivatives: $$\grad f(\x)=\left(\frac{\partial f}{\partial x_1}(\x),\ \dots,\ \frac{\partial f}{\partial x_n}(\x)\right)^\top\in\R^n.$$ The symbol $\grad$ is read "nabla" or "grad". If all partial derivatives exist and are continuous, we say $f$ is continuously differentiable and write $f\in C^1$.

Slope in an arbitrary direction

Partial derivatives give slopes along the coordinate axes. What about along some other direction $\d$, say diagonally? Walk along the straight line $t\mapsto\x+t\d$ and measure how fast $f$ changes as $t$ starts increasing from 0. That rate is the directional derivative: $$D_{\d}f(\x)=\lim_{t\to0^+}\frac{f(\x+t\d)-f(\x)}{t}.$$ Here is the fact that makes the gradient so useful: you never need to compute this limit, because the gradient already knows the answer.

If $f\in C^1$, then for every direction $\d$: $\quad D_{\d}f(\x)=\grad f(\x)^\top\d.$

In particular $\d$ is a descent direction (moving along it decreases $f$ at first) whenever $\grad f(\x)^\top\d<0$.

Why it's true, in one line: let $g(t)=f(\x+t\d)$. Each coordinate $x_i+td_i$ moves at speed $d_i$, so by the chain rule $g'(0)=\sum_i\frac{\partial f}{\partial x_i}(\x)\,d_i=\grad f(\x)^\top\d$.

If $\grad f(\x)\ne\0$, then among all unit vectors $\d$ ($\norm{\d}=1$), the slope $D_{\d}f(\x)$ is largest for $\d=\grad f(\x)/\norm{\grad f(\x)}$, where it equals $\norm{\grad f(\x)}$, and smallest for $\d=-\grad f(\x)/\norm{\grad f(\x)}$, where it equals $-\norm{\grad f(\x)}$.

Proof. By Cauchy–Schwarz (Part 0a), $-\norm{\grad f}\norm{\d}\le\grad f^\top\d\le\norm{\grad f}\norm{\d}$, and $\norm{\d}=1$. Equality on the right needs $\d$ to be a positive multiple of $\grad f$; on the left, a negative multiple. With length 1, that pins $\d$ down exactly. $\blacksquare$

This is the whole reason gradient descent steps along $-\grad f$: locally, it is the steepest way down. ("Locally" matters. Part 5 will show that steepest is not always fastest overall.)

One more picture worth keeping: the gradient is perpendicular to the level set through $\x$. Moving along a contour line keeps $f$ constant, so the slope along it is zero, so $\grad f^\top\d=0$ for the tangent direction $\d$.

Try it

Drag the black point anywhere on the map. The blue arrow is $\grad f$ (uphill); the green arrow is $-\grad f$ (steepest downhill). Turn the dial to choose a direction $\d$ and read its slope $\grad f^\top\d$. Find the direction with the most negative slope, and check it matches the green arrow.

Let $f(x_1,x_2)=x_1^2+4x_2^2$ at $\x=(1,1)^\top$. Find $\grad f(\x)$, the slope along $\d=(1,-1)^\top/\sqrt2$, and the steepest descent direction.

  1. $\frac{\partial f}{\partial x_1}=2x_1$, $\frac{\partial f}{\partial x_2}=8x_2$. At $(1,1)$: $\grad f=(2,8)^\top$.

    Differentiate in one variable at a time, holding the other fixed.

  2. $D_{\d}f=\grad f^\top\d=(2\cdot1+8\cdot(-1))/\sqrt2=-6/\sqrt2\approx-4.24$.

    The theorem lets us replace a limit by a dot product. Negative means $\d$ goes downhill.

  3. Steepest descent: $-\grad f/\norm{\grad f}=-(2,8)^\top/\sqrt{68}\approx(-0.243,-0.970)^\top$, with slope $-\sqrt{68}\approx-8.25$.

    Any unit direction has slope at least $-\norm{\grad f}$; this one achieves it.

  4. Note how the steepest direction points mostly along $x_2$, not towards the minimizer $\0$ (which is in direction $(-1,-1)/\sqrt{2}$).

    On a stretched bowl, "steepest" and "towards the bottom" disagree. This is the seed of gradient descent's zig-zag in Part 5.

Let $f(x_1,x_2)=x_1^2x_2+3x_1-x_2$. Compute both partial derivatives at $(1,2)$.

$\partial f/\partial x_1=2x_1x_2+3$ (treat $x_2$ as a constant). Now do $\partial f/\partial x_2$ the same way.

$\partial f/\partial x_1=2x_1x_2+3=2(1)(2)+3=7$. $\ \partial f/\partial x_2=x_1^2-1=1-1=0$. So $\grad f(1,2)=(7,0)^\top$: at this point, $f$ is flat in the $x_2$ direction.

For $f(\x)=x_1^2+x_2^2$ at $\x=(3,4)^\top$, compute the directional derivative $D_{\d}f(\x)=\grad f(\x)^\top\d$ along the unit vector $\d=(-0.6,\,0.8)^\top$.

$\grad f=(2x_1,2x_2)^\top=(6,8)^\top$.

$(6)(-0.6)+(8)(0.8)=-3.6+6.4=2.8$.

At some point, $\grad f(\x)=(3,-4)^\top$. What is the most negative directional derivative over all unit vectors $\d$?

The steepest descent slope is $-\norm{\grad f(\x)}$.

$-\norm{(3,-4)}=-\sqrt{9+16}=-5$, attained at $\d=(-3,4)^\top/5$.

  • Hold the other variables fixed when taking a partial derivative
  • Use $\grad f^\top\d$ to test whether a direction goes downhill
  • Normalize $\d$ to length 1 before comparing slopes of different directions
  • Remember the gradient is perpendicular to contour lines
  • Thinking $-\grad f$ points at the minimizer (it points steepest-downhill here, which is different)
  • Comparing $\grad f^\top\d$ for directions of different lengths (longer $\d$ inflates the number)
  • Forgetting the product rule in terms like $x_1^2x_2$
  1. $\grad f(\x)$ collects the partial derivatives; the slope along any direction is $D_{\d}f(\x)=\grad f(\x)^\top\d$.
  2. $\d$ is a descent direction whenever $\grad f(\x)^\top\d<0$ (and it is not one if $\grad f(\x)^\top\d>0$).
  3. By Cauchy–Schwarz, $-\grad f(\x)/\norm{\grad f(\x)}$ is the steepest descent direction, with slope $-\norm{\grad f(\x)}$.

At a point, $\grad f=(2,-1)^\top$. Which direction is a descent direction?

$(1,0)^\top$
$\grad f^\top\d=2>0$: that goes uphill.
$(-1,-1)^\top$
$\grad f^\top\d=-2+1=-1<0$, so $f$ decreases when you start moving along it.
$(1,2)^\top$
$\grad f^\top\d=2-2=0$: it runs along the contour, neither up nor down at first.

The gradient at a point is…

tangent to the contour line through that point
Along the contour, $f$ doesn't change, so the slope there is zero, and the slope is $\grad f^\top\d$. What does that say about the angle?
perpendicular to the contour line through that point
Slope along the contour is 0, so $\grad f^\top\d=0$ for the tangent $\d$: perpendicular.
always pointing towards the global maximum
It points steepest uphill locally, which need not be towards any maximum.

Why does the steepest-ascent theorem need $\grad f(\x)\ne\0$?

Because the formula divides by $\norm{\grad f(\x)}$, and when the gradient is zero every direction has slope 0
At a flat point there's no steepest direction: all slopes are $\0^\top\d=0$.
Because the gradient is undefined when it's zero
The zero vector is a perfectly good gradient; it just means "flat here".
Because Cauchy–Schwarz fails for the zero vector
Cauchy–Schwarz holds for all vectors, including $\0$ (both sides are 0).

For $f(x_1,x_2)=e^{x_1}\sin x_2$, what is $\frac{\partial f}{\partial x_2}$?

$e^{x_1}\sin x_2$
That's $\partial f/\partial x_1$, since $e^{x_1}$ differentiates to itself.
$e^{x_1}\cos x_2$
Treat $e^{x_1}$ as a constant and differentiate $\sin x_2$.
$e^{x_1}\cos x_2+e^{x_1}\sin x_2$
No product rule is needed: only one factor depends on $x_2$.

Matrices, eigenvalues and positive definiteness

A symmetric matrix stretches space along perpendicular directions (its eigenvectors) by amounts (its eigenvalues), and positive definite means "stretches every direction by a positive amount".

Curvature of a function in many variables is a matrix (the Hessian). Whether a critical point is a minimum, and how fast algorithms converge, is read straight off its eigenvalues.

A rubber sheet pulled by two people standing at right angles: each pulls along their own line, and the eigenvalues say how hard each one pulls.

A matrix is a machine that moves vectors

A $2\times2$ matrix $A=\begin{pmatrix}a&b\\c&d\end{pmatrix}$ turns a vector $\v=(v_1,v_2)^\top$ into $$A\v=\begin{pmatrix}av_1+bv_2\\ cv_1+dv_2\end{pmatrix}.$$ Each entry of $A\v$ is a dot product of a row of $A$ with $\v$. The transpose $A^\top$ swaps rows and columns. $A$ is symmetric if $A^\top=A$, i.e. $b=c$. Every Hessian in this course is symmetric, so symmetric matrices are our focus.

The identity $I$ has 1s on the diagonal and 0s elsewhere: $I\v=\v$. The inverse $A^{-1}$, when it exists, undoes $A$: $A^{-1}A=I$. For $2\times2$: $A^{-1}=\frac{1}{ad-bc}\begin{pmatrix}d&-b\\-c&a\end{pmatrix}$, which exists when the determinant $\det A=ad-bc\ne0$.

Eigenvectors: directions that don't turn

Most vectors get rotated by $A$. A few special directions only get stretched (or flipped). If $$A\u=\lambda\u,\qquad \u\ne\0,$$ then $\u$ is an eigenvector and the number $\lambda$ ("lambda") is its eigenvalue: the stretch factor. To find eigenvalues, solve $\det(A-\lambda I)=0$. For a symmetric $2\times 2$ matrix $\begin{pmatrix}a&b\\b&d\end{pmatrix}$: $$\lambda^2-(a+d)\lambda+(ad-b^2)=0.$$ Useful shortcuts: the eigenvalues add up to the trace $a+d$ and multiply to the determinant $ad-b^2$.

If $A$ is a real symmetric $n\times n$ matrix, then all its eigenvalues are real, and it has $n$ eigenvectors $\u_1,\dots,\u_n$ that are mutually perpendicular and of length 1 (an orthonormal basis). Equivalently $A=U\Lambda U^\top$, where the columns of $U$ are the $\u_i$ and $\Lambda$ is the diagonal matrix of eigenvalues.

What this buys us: write any vector in the eigenvector "coordinates", $\v=c_1\u_1+\dots+c_n\u_n$. Then $$\v^\top A\v=\lambda_1c_1^2+\lambda_2c_2^2+\dots+\lambda_nc_n^2.$$ The expression $\v^\top A\v$ is called the quadratic form of $A$. In eigen-coordinates it's just a weighted sum of squares, with the eigenvalues as weights. That one line explains everything below.

A symmetric matrix $A$ is

  • positive definite ($A\succ0$) if $\v^\top A\v>0$ for every $\v\ne\0$;
  • positive semidefinite ($A\succeq0$) if $\v^\top A\v\ge0$ for every $\v$;
  • negative (semi)definite if $-A$ is positive (semi)definite;
  • indefinite if $\v^\top A\v$ is positive for some $\v$ and negative for others.

With eigenvalues $\lambda_1\le\dots\le\lambda_n$: $\ A\succ0\iff\lambda_1>0$; $\ A\succeq0\iff\lambda_1\ge0$; $\ A$ indefinite $\iff\lambda_1<0<\lambda_n$.

Rayleigh bounds: for every $\v$, $\ \lambda_1\norm{\v}^2\le\v^\top A\v\le\lambda_n\norm{\v}^2$, with equality at $\v=\u_1$ and $\v=\u_n$ respectively.

Both follow from the weighted-sum-of-squares formula: if every weight $\lambda_i$ is positive, a sum of $\lambda_i c_i^2$ is positive unless all $c_i=0$. And the sum lies between $\lambda_1\sum c_i^2$ and $\lambda_n\sum c_i^2$, where $\sum c_i^2=\norm{\v}^2$.

A quick hand test (Sylvester's criterion). A symmetric matrix is positive definite if and only if all its leading principal minors are positive: the top-left $1\times1$ determinant, the top-left $2\times2$ determinant, and so on. For $2\times2$: $a>0$ and $ad-b^2>0$.

The condition number. For $A\succ0$, $\kappa(A)=\lambda_{\max}/\lambda_{\min}\ge1$. The level sets $\{\v:\v^\top A\v=1\}$ are ellipses with half-axes $1/\sqrt{\lambda_i}$ along $\u_i$. So $\kappa$ says how stretched they are: the ratio of longest to shortest axis is $\sqrt\kappa$. A large $\kappa$ means a long narrow valley, and that is what slows gradient descent down.

Try it

Change the entries of the symmetric matrix $A=\begin{pmatrix}a&b\\b&d\end{pmatrix}$. The plot shows the level curves of $\v^\top A\v$, the eigenvector directions, and the verdict. Find settings that make $A$: positive definite with $\kappa=10$; semidefinite but not definite; indefinite.

Classify $A=\begin{pmatrix}2&1\\1&2\end{pmatrix}$ and find its condition number and eigenvectors.

  1. Characteristic equation: $\lambda^2-4\lambda+(4-1)=\lambda^2-4\lambda+3=0$, so $\lambda=1$ or $3$.

    Trace $=4$ and determinant $=3$ give the sum and product of the eigenvalues directly.

  2. Both eigenvalues are positive, so $A\succ0$. Sylvester agrees: $2>0$ and $\det A=3>0$.

    Two independent tests reaching the same verdict is a good habit on exams.

  3. For $\lambda=3$: $(A-3I)\u=\0$ gives $-u_1+u_2=0$, so $\u_2=(1,1)^\top/\sqrt2$. For $\lambda=1$: $u_1+u_2=0$, so $\u_1=(1,-1)^\top/\sqrt2$.

    The eigenvectors are perpendicular, as the spectral theorem promises.

  4. $\kappa=3/1=3$. The level curves of $\v^\top A\v$ are ellipses tilted at 45°, $\sqrt3\approx1.73$ times longer along $(1,-1)$ than along $(1,1)$.

    Half-axis lengths are $1/\sqrt{\lambda}$: the small eigenvalue gives the long axis.

Find the eigenvalues of $A=\begin{pmatrix}5&2\\2&2\end{pmatrix}$.

Trace $=7$, determinant $=10-4=6$. Which two numbers add to 7 and multiply to 6?

$\lambda^2-7\lambda+6=(\lambda-1)(\lambda-6)=0$, so $\lambda=1$ and $6$.

Classify $A=\begin{pmatrix}1&3\\3&1\end{pmatrix}$.

The determinant is $1-9=-8$. If the product of the two eigenvalues is negative, what are their signs?

Eigenvalues $1\pm3$, i.e. $4$ and $-2$: one positive, one negative, so indefinite. For example $\v=(1,1)$ gives $\v^\top A\v=8>0$ and $\v=(1,-1)$ gives $-4<0$.

What is the condition number of the matrix in Problem 1?

$\kappa=\lambda_{\max}/\lambda_{\min}$.

$\kappa=6/1=6$.

For which values of $t$ is $\begin{pmatrix}1&2\\2&t\end{pmatrix}$ positive semidefinite? Enter the smallest such $t$.

For $2\times2$ symmetric with $a=1>0$: semidefinite needs $\det\ge0$.

$\det=t-4\ge0$ (together with diagonal entries $\ge 0$) gives $t\ge4$. At $t=4$ the eigenvalues are $0$ and $5$: semidefinite but not definite.

  • Use trace and determinant to get 2×2 eigenvalues fast
  • Check definiteness two ways (eigenvalues and Sylvester) when it matters
  • Think "weighted sum of squares $\sum\lambda_ic_i^2$" whenever you see $\v^\top A\v$
  • Using Sylvester's leading-minors test (with $\ge0$) for semidefiniteness (for that, all principal minors, not only leading ones, must be $\ge0$)
  • Deciding definiteness from the diagonal entries alone ($\begin{pmatrix}1&3\\3&1\end{pmatrix}$ has a positive diagonal and is indefinite)
  • Mixing up which eigenvalue gives the long axis (the small one)
  1. A symmetric matrix has real eigenvalues and perpendicular eigenvectors; in those coordinates $\v^\top A\v=\sum_i\lambda_ic_i^2$.
  2. Positive definite $\iff$ all eigenvalues positive; Rayleigh: $\lambda_{\min}\norm{\v}^2\le\v^\top A\v\le\lambda_{\max}\norm{\v}^2$.
  3. $\kappa=\lambda_{\max}/\lambda_{\min}$ measures how stretched the level ellipses are; large $\kappa$ means a narrow valley.

A symmetric $3\times3$ matrix has eigenvalues $-1, 0, 5$. It is…

positive semidefinite
Semidefinite needs every eigenvalue $\ge0$; one of them is $-1$.
indefinite
One negative and one positive eigenvalue: $\v^\top A\v$ takes both signs.
singular, so we can't tell
A zero eigenvalue makes it singular, but the signs of the others still decide definiteness.

If $A\succ0$ with $\lambda_{\min}=2$ and $\lambda_{\max}=8$, which inequality holds for all $\v$?

$2\norm{\v}^2\le\v^\top A\v\le8\norm{\v}^2$
That's the Rayleigh bound.
$\v^\top A\v\le2\norm{\v}^2$
Along the top eigenvector, $\v^\top A\v=8\norm{\v}^2$, which breaks this.
$\v^\top A\v=5\norm{\v}^2$
It's only constant if all eigenvalues are equal.

The level sets of $\v^\top A\v$ for $A\succ0$ are ellipses with axis ratio 10. What is $\kappa(A)$?

10
Axis lengths go like $1/\sqrt{\lambda}$, not $1/\lambda$. Redo the ratio.
100
Axis ratio $=\sqrt{\lambda_{\max}/\lambda_{\min}}=10$, so $\kappa=100$.
$\sqrt{10}$
You took the square root the wrong way round.

Which matrix is positive definite?

$\begin{pmatrix}1&2\\2&1\end{pmatrix}$
$\det=1-4<0$: eigenvalues of opposite sign.
$\begin{pmatrix}0&0\\0&3\end{pmatrix}$
Eigenvalue 0 means only semidefinite: $\v=(1,0)$ gives $\v^\top A\v=0$.
$\begin{pmatrix}3&1\\1&1\end{pmatrix}$
$3>0$ and $\det=3-1=2>0$: Sylvester says positive definite.

Curvature: the Hessian and Taylor's theorem

Taylor's theorem says that near any point, a smooth function looks like a quadratic bowl (or saddle) built from its gradient and its Hessian.

Every theorem in the next parts, from "when is a point a minimum" to "how fast does Newton's method converge", is proved by writing down a Taylor expansion.

A curved road seen through a magnifying glass: zoom in enough and it looks straight (first order); zoom a bit less and you see a gentle bend (second order).

Second derivatives in many variables

In one variable, $f''(x)$ measures how the slope changes: $f''>0$ means the graph curves upward like a cup. In $n$ variables there are $n^2$ second derivatives $\frac{\partial^2 f}{\partial x_i\partial x_j}$ (differentiate with respect to $x_j$, then $x_i$). We arrange them in a matrix.

For $f$ with continuous second partial derivatives (written $f\in C^2$), the Hessian at $\x$ is the $n\times n$ matrix $$\big[\hess f(\x)\big]_{ij}=\frac{\partial^2 f}{\partial x_i\,\partial x_j}(\x).$$

Example: $f=x_1^2+3x_1x_2+5x_2^2$. Then $\grad f=(2x_1+3x_2,\ 3x_1+10x_2)^\top$ and $$\hess f=\begin{pmatrix}2&3\\3&10\end{pmatrix}.$$ Notice the matrix is symmetric: $\partial^2f/\partial x_1\partial x_2=\partial^2f/\partial x_2\partial x_1=3$. That's no accident.

If the mixed partial derivatives $\frac{\partial^2 f}{\partial x\partial y}$ and $\frac{\partial^2 f}{\partial y\partial x}$ exist near a point and are continuous at it, they are equal there. Hence for every $f\in C^2$, the Hessian is symmetric, and the spectral theorem applies to it.

Go deeper: why continuity matters (Peano's counterexample)

Take $f(x,y)=\dfrac{xy(x^2-y^2)}{x^2+y^2}$ for $(x,y)\ne(0,0)$ and $f(0,0)=0$. A direct calculation gives $\frac{\partial f}{\partial x}(0,y)=-y$ and $\frac{\partial f}{\partial y}(x,0)=x$. Differentiating again at the origin: $\frac{\partial^2 f}{\partial y\,\partial x}(0,0)=-1$ but $\frac{\partial^2 f}{\partial x\,\partial y}(0,0)=+1$. The mixed partials exist but disagree, because they are not continuous at the origin (in polar form $f=\frac{r^2}{4}\sin4\theta$, and the mixed partial keeps oscillating with $\theta$ as $r\to0$). So "the Hessian is symmetric" is a theorem with a hypothesis, $f\in C^2$, and this is the example showing why the hypothesis is needed.

Proof idea of Schwarz's theorem. Look at the second difference $\Delta(h,k)=f(h,k)-f(h,0)-f(0,k)+f(0,0)$. Applying the one-variable mean value theorem twice, first in $x$ then in $y$, gives $\Delta/(hk)=\partial_{yx}f(\xi,\eta)$ at some point near the origin; doing it in the other order gives $\Delta/(hk)=\partial_{xy}f(\tilde\xi,\tilde\eta)$. Let $h,k\to0$ and use continuity: both sides tend to the values at the origin, which must therefore be equal.

Taylor's theorem: the local quadratic picture

The trick that turns many-variable calculus into one-variable calculus: walk along a straight line. Fix $\x$ and a step $\p$, and let $g(t)=f(\x+t\p)$. The chain rule gives $$g'(t)=\grad f(\x+t\p)^\top\p,\qquad g''(t)=\p^\top\hess f(\x+t\p)\,\p.$$ Now apply the school Taylor formula to $g$ between $t=0$ and $t=1$.

  1. Lagrange form (exact, with an unknown middle point): for some $\theta\in(0,1)$, $$f(\x+\p)=f(\x)+\grad f(\x)^\top\p+\tfrac12\,\p^\top\hess f(\x+\theta\p)\,\p.$$
  2. Asymptotic form (approximate, for small $\p$): $$f(\x+\p)=f(\x)+\grad f(\x)^\top\p+\tfrac12\,\p^\top\hess f(\x)\,\p+o(\norm{\p}^2).$$
  3. Integral form (exact, first order, $f\in C^1$ suffices): $$f(\x+\p)=f(\x)+\int_0^1\grad f(\x+t\p)^\top\p\,dt.$$

The symbol $o(\norm{\p}^2)$ ("little-o") means "an error that shrinks faster than $\norm{\p}^2$": divide it by $\norm{\p}^2$ and the result goes to 0 as $\p\to\0$. So for small steps, the quadratic part dominates the error.

Which form when? The book's advice, worth memorizing: the Lagrange form proves necessary conditions for a minimum (Part 3); the asymptotic form proves sufficient conditions (Part 3); the integral form proves global inequalities like the descent lemma (Part 6).

Quadratics are their own Taylor expansion

The most important functions in this course are quadratics $$f(\x)=\tfrac12\x^\top A\x-\b^\top\x+c,\qquad A\text{ symmetric}.$$ For them, $\grad f(\x)=A\x-\b$ and $\hess f(\x)=A$ everywhere, and the second-order Taylor formula is exact with no error term. Our fit-the-line loss is such a quadratic: expanding the squares gives $L(w,c)=\frac12\z^\top A\z-\b^\top\z+38$ with $\z=(w,c)^\top$, $$A=\begin{pmatrix}28&12\\12&6\end{pmatrix},\qquad \b=\begin{pmatrix}46\\20\end{pmatrix}.$$ Its eigenvalues are about $0.72$ and $33.3$, so $\kappa\approx46$: that's why the valley in Part 0a looked so long and thin.

Try it

Pick a function and drag the expansion point $x_0$ (or use the slider). Compare the tangent line (first order) with the Taylor parabola (second order). Zoom in with the window slider: the parabola hugs the curve over a wider range than the line.

Write the second-order Taylor expansion of $f(x_1,x_2)=e^{x_1}+x_1x_2^2$ around $\x=(0,1)^\top$, and use it to estimate $f(0.1,\,0.9)$.

  1. $f(0,1)=1+0=1$.

    Start with the value at the expansion point.

  2. $\grad f=(e^{x_1}+x_2^2,\ 2x_1x_2)^\top$, so $\grad f(0,1)=(2,0)^\top$.

    Each partial derivative treats the other variable as a constant.

  3. $\hess f=\begin{pmatrix}e^{x_1}&2x_2\\2x_2&2x_1\end{pmatrix}$, so $\hess f(0,1)=\begin{pmatrix}1&2\\2&0\end{pmatrix}$.

    Differentiate each gradient entry again. Symmetric, as Schwarz promised.

  4. With $\p=(0.1,-0.1)^\top$: $\grad f^\top\p=0.2$ and $\p^\top\hess f\,\p=0.01-0.04+0=-0.03$. So $f\approx1+0.2-0.015=1.185$.

    $\p^\top H\p=H_{11}p_1^2+2H_{12}p_1p_2+H_{22}p_2^2=0.01+2(2)(0.1)(-0.1)+0$.

  5. The true value is $e^{0.1}+0.1(0.81)\approx1.10517+0.081=1.18617$. The error, about 0.001, is third order in the step size.

    Checking against the exact value builds trust that the quadratic model is good for small steps.

Compute the Hessian of $f(x_1,x_2)=x_1^3-x_1x_2+x_2^2$ at $(1,0)$.

$\grad f=(3x_1^2-x_2,\ -x_1+2x_2)^\top$. Differentiate each entry again with respect to $x_1$ and $x_2$.

$\hess f=\begin{pmatrix}6x_1&-1\\-1&2\end{pmatrix}$; at $(1,0)$: $H_{11}=6$, $H_{12}=-1$, $H_{22}=2$.

Use the second-order Taylor expansion of $f(x)=x^2+x$ around $x_0=1$ to estimate $f(1.01)$. (For a quadratic, the estimate is exact.)

$f(1)=2$, $f'(1)=3$, $f''=2$. Use $f(1+h)\approx f(1)+f'(1)h+\tfrac12f''h^2$ with $h=0.01$.

$2+3(0.01)+\tfrac12(2)(0.0001)=2+0.03+0.0001=2.0301$. Exact: $1.0201+1.01=2.0301$.

For the quadratic $f(\x)=\tfrac12\x^\top A\x-\b^\top\x$ with $A=\begin{pmatrix}2&0\\0&4\end{pmatrix}$, $\b=(-4,6)^\top$, compute $\grad f(\0)$.

$\grad f(\x)=A\x-\b$. At $\x=\0$ this is just $-\b$.

$\grad f(\0)=-\b=(4,-6)^\top$.

  • Reduce many-variable questions to one variable via $g(t)=f(\x+t\p)$
  • Remember $\grad(\tfrac12\x^\top A\x-\b^\top\x)=A\x-\b$ and $\hess=A$ for symmetric $A$
  • Check your Hessian is symmetric; if not, you've made an arithmetic slip
  • Keep the $\tfrac12$ in front of the quadratic term
  • Dropping the factor $\tfrac12$ in $\tfrac12\p^\top H\p$
  • Forgetting the factor 2 on the off-diagonal term: $\p^\top H\p=H_{11}p_1^2+2H_{12}p_1p_2+H_{22}p_2^2$
  • Writing $\grad(\x^\top A\x)=A\x$ (it's $2A\x$ for symmetric $A$)
  • Using the asymptotic form as if it were exact for large steps
  1. The Hessian collects second partial derivatives; it's symmetric whenever $f\in C^2$ (Schwarz).
  2. Taylor: $f(\x+\p)\approx f(\x)+\grad f(\x)^\top\p+\frac12\p^\top\hess f(\x)\p$, with exact Lagrange and integral versions.
  3. Quadratics $\frac12\x^\top A\x-\b^\top\x$ have gradient $A\x-\b$ and Hessian $A$; their Taylor expansion is exact.

For $f(\x)=\tfrac12\x^\top A\x-\b^\top\x$ with $A$ symmetric, the gradient is:

$A\x$
You've dropped the linear term's contribution, $-\b$.
$A\x-\b$
The $\frac12$ cancels the 2 from differentiating the quadratic form, and $-\b^\top\x$ contributes $-\b$.
$2A\x-\b$
That would be the gradient of $\x^\top A\x-\b^\top\x$, without the $\frac12$.

Which form of Taylor's theorem is exact but involves an unknown point $\x+\theta\p$?

Asymptotic (with $o(\norm{\p}^2)$)
That form is approximate for small $\p$; it uses the Hessian at $\x$ itself.
Lagrange form
It is exact, but the Hessian is evaluated at an unknown $\x+\theta\p$ with $\theta\in(0,1)$.
None of them is exact
Both the Lagrange and the integral forms are exact.

Why is the Hessian of every $C^2$ function symmetric?

Because all matrices that come from functions are symmetric
Not in general: Peano's example has unequal mixed partials. Something about $C^2$ is needed.
By Schwarz's theorem: continuous mixed partials are equal
$C^2$ means the second partials are continuous, which is exactly Schwarz's hypothesis.
Because $\frac{\partial}{\partial x}$ and $\frac{\partial}{\partial y}$ always commute
They don't always: Peano's counterexample shows they can fail to.

At $\x$, $\grad f(\x)=(0,2)^\top$. Moving along $\d=(5,0)^\top$, $f$ initially…

decreases
$\grad f^\top\d=0\cdot5+2\cdot0=0$, so the slope is zero, not negative.
neither increases nor decreases to first order
The direction is perpendicular to the gradient, i.e. along the contour.
increases
Compute $\grad f^\top\d$: it's 0.

A symmetric $2\times2$ matrix has trace 5 and determinant −6. It is…

positive definite
A negative determinant means the eigenvalues multiply to a negative number.
indefinite
Eigenvalues multiply to −6, so one is positive and one negative (here 6 and −1).
negative definite
Two negative eigenvalues would multiply to a positive number.

Using $\hess f(\x)=\begin{pmatrix}4&1\\1&2\end{pmatrix}$, what is $\p^\top\hess f(\x)\p$ for $\p=(1,-1)^\top$?

6
You forgot the off-diagonal terms: $\p^\top H\p$ includes $2H_{12}p_1p_2$.
4
$4(1)^2+2(1)(1)(-1)+2(-1)^2=4-2+2=4$.
8
The cross term has a minus sign here because $p_1p_2=-1$.

Cauchy–Schwarz is the key step in proving…

that the Hessian is symmetric
That's Schwarz's theorem on mixed partials, which is a different Schwarz result.
that $-\grad f/\norm{\grad f}$ is the steepest descent direction
It bounds $\grad f^\top\d$ below by $-\norm{\grad f}$ for unit $\d$.
that eigenvalues of symmetric matrices are real
That's part of the spectral theorem, proved differently.

Lecture 1 of the course: what an optimization problem is, the difference between a local and a global answer, the 1945 diet problem that helped start the field, and why you must classify a problem before choosing an algorithm.

You need: Part 0a (notation, min vs argmin, level sets). Part 0b helps for Chapter 1.3.

The anatomy of an optimization problem

An optimization problem has three parts (what you can choose, what you want to make small, and what rules you must obey), and its answer can be the best overall or only the best nearby.

Most real tasks only become solvable once you write them in this form. And knowing whether an answer is "local" or "global" is the difference between a good solution and the best one.

Planning a trip: you choose the route (decision variables), you want the shortest travel time (objective), and you must arrive before the shop closes (constraint). The best route through your own neighbourhood isn't necessarily the best route overall.

Every problem in this course has the same three ingredients:

  • Decision variables $\x=(x_1,\dots,x_n)^\top$: the numbers you're free to choose.
  • Objective function $f(\x)$: a single number measuring how bad a choice is (a cost, an error, an energy). We minimize it. To maximize a profit $P$, minimize $-P$: same problem.
  • Constraints: rules $g_i(\x)\le0$ and $h_j(\x)=0$ that a choice must satisfy. The choices that obey every rule form the feasible set $\mathcal F$.

Two ways a problem can go wrong before you even start: it can be infeasible ($\mathcal F=\emptyset$: no choice obeys all the rules), or unbounded (you can make $f$ as negative as you like, so $\inf f=-\infty$ and there's no minimizer). And even with a finite infimum, it may never be reached, as with $\min_x e^x$; again there is no minimizer.

Formulating real tasks

Fitting a model (ridge regression). You have a data matrix $A$ (one row per example) and targets $\b$. Choose model weights $\x$ to make predictions $A\x$ close to $\b$, without letting the weights grow wild: $$\min_{\x\in\R^n}\ \tfrac12\norm{A\x-\b}^2+\tfrac\lambda2\norm{\x}^2,\qquad\lambda>0.$$ Unconstrained. Expanding gives the quadratic $\tfrac12\x^\top(A^\top A+\lambda I)\x-(A^\top\b)^\top\x+\tfrac12\norm{\b}^2$, whose matrix $A^\top A+\lambda I$ is positive definite. Our fit-the-line example is the case $\lambda=0$, up to a factor of 2 (its matrix $A$ has rows $(x_i,1)$).

Investing (Markowitz portfolio). Split your money into fractions $w_1,\dots,w_n$ across $n$ assets with expected returns $\boldsymbol\mu$ and covariance matrix $\Sigma$. Minimize risk $\frac12\mathbf w^\top\Sigma\mathbf w$ subject to an expected return of at least $r$ ($\boldsymbol\mu^\top\mathbf w\ge r$), no short-selling ($w_i\ge0$), and investing everything ($\sum_iw_i=1$). Constrained.

Global and local minimizers

A ball of radius $\varepsilon$ around $\x^\star$ is $B(\x^\star,\varepsilon)=\{\x:\norm{\x-\x^\star}<\varepsilon\}$: every point closer than $\varepsilon$.

  • $\x^\star\in\mathcal F$ is a global minimizer if $f(\x^\star)\le f(\x)$ for every $\x\in\mathcal F$.
  • $\x^\star\in\mathcal F$ is a local minimizer if there is some $\varepsilon>0$ with $f(\x^\star)\le f(\x)$ for every feasible $\x$ in the ball $B(\x^\star,\varepsilon)$.
  • Replace $\le$ by $<$ (for $\x\ne\x^\star$) to get strict global/local minimizers.

The optimal value is $p^\star=\inf_{\x\in\mathcal F}f(\x)$.

Every global minimizer is a local one, but not the other way round. A local minimizer only has to beat its neighbours. Think of a mountain lake: it's the lowest point around, but the sea is lower.

The algorithms in this course look only at local information (the value, slope and curvature where they stand). So by themselves they can only promise local answers. Part 4 (convexity) identifies the problems where local automatically means global.

Try it

This curve is $f(x)=x^4+a\,x^3+b\,x^2+c\,x$. Move the sliders. Local minima are marked in green, local maxima in orange, and the global minimum gets a ring. Find settings with two local minima, then make the "wrong" one become global.

Find all local and global minimizers of $f(x)=x^4-2x^2$ on $\R$, and of the same $f$ on the interval $[0,\tfrac12]$.

  1. $f(x)=x^2(x^2-2)$, and $f'(x)=4x^3-4x=4x(x-1)(x+1)$, which is zero at $x=-1,0,1$.

    A smooth function on an open set can only have a local minimum where its slope is zero (Part 3 proves this). So these are the candidates.

  2. $f(\pm1)=-1$ and $f(0)=0$. Near 0, $f(x)\approx-2x^2<0=f(0)$, so 0 is a local maximum. Both $x=\pm1$ are local minima, and since $f(x)=(x^2-1)^2-1\ge-1$ everywhere, they're both global.

    Rewriting as a square plus a constant proves "global" in one line: a square is never negative.

  3. On $[0,\tfrac12]$: $f'(x)=4x(x^2-1)<0$ for $0\lt x\le\tfrac12$, so $f$ is decreasing there. The minimum is at the right end, $x^\star=\tfrac12$, with $f=\tfrac1{16}-\tfrac12=-\tfrac7{16}$. So $x=\tfrac12$ is the only local (and global) minimizer on $[0,\tfrac12]$; the left end $x=0$ is a local maximum there.

    With constraints, a minimizer can sit on the boundary where $f'\ne0$. The "slope zero" rule is only for interior points.

Classify $\min\ x_1-x_2$ subject to $x_1\ge0$, $x_2\ge0$.

Can you make $x_1-x_2$ as negative as you like while staying feasible? Try $x_1=0$ and a huge $x_2$.

Take $x_1=0$, $x_2=t$: feasible for every $t\ge0$, and the objective is $-t\to-\infty$. The problem is unbounded: $p^\star=-\infty$, no minimizer.

What is the minimum value of $x^4-2x^2$ over $x\in[0,\tfrac12]$?

The function is decreasing on this interval (check the sign of $f'$), so evaluate at the right end.

$f(\tfrac12)=\tfrac1{16}-\tfrac24=-\tfrac7{16}=-0.4375$.

How many local minimizers does $f(x)=\cos x+\tfrac{x^2}{20}$ have on the interval $(-4,4)$? (Use the explorer intuition: a gentle bowl with ripples.)

$f'(x)=-\sin x+x/10$. Near $x=0$, $f''(0)=-1+0.1<0$: is 0 a minimum or a maximum? Where else might the slope vanish?

$f'(x)=0$ means $\sin x=x/10$, which has solutions $x=0$ and $x\approx\pm2.85$ in $(-4,4)$. At 0, $f''(0)=-0.9<0$: a local maximum. At $x\approx\pm2.85$, $f''=-\cos x+0.1\approx0.96+0.1>0$: local minima. So there are 2.

  • Name the decision variables, objective and constraints explicitly before anything else
  • Turn "maximize $P$" into "minimize $-P$"
  • Ask "is a local answer good enough, or do I need the global one?"
  • Check the boundary when there are constraints
  • Calling a point global just because the algorithm stopped there
  • Forgetting that an unconstrained linear objective is always unbounded (unless it's constant)
  • Applying "$f'(x)=0$" at a boundary point of the feasible set
  1. A problem = decision variables + objective + constraints; the feasible set $\mathcal F$ is everything that obeys the rules.
  2. Global minimizers beat every feasible point; local ones only beat their neighbours within some radius $\varepsilon$.
  3. Problems can also be infeasible ($\mathcal F=\emptyset$) or unbounded ($p^\star=-\infty$).

Which statement is always true?

Every local minimizer is a global minimizer
Think of $x^4-2x^2+x$: it has a local minimizer that's higher than the other one.
Every global minimizer is a local minimizer
If you beat every feasible point, you certainly beat those in a small ball.
Every function has at least one local minimizer
$f(x)=x$ on $\R$ has none: you can always step left to go lower.

To maximize profit $P(\x)$ with a minimization solver, you…

minimize $1/P(\x)$
That breaks when $P\le0$ and changes the problem's shape. There's a cleaner trick.
minimize $-P(\x)$
The point that makes $-P$ smallest makes $P$ largest, and the optimal values are negatives of each other.
minimize $P(\x)^2$
That finds where $P$ is closest to zero, not where it's largest.

Ridge regression $\min\frac12\norm{A\x-\b}^2+\frac\lambda2\norm{\x}^2$ is…

an unconstrained problem with a positive definite quadratic objective
There are no constraints, and $A^\top A+\lambda I\succ0$ for $\lambda>0$.
a constrained problem because of the penalty $\norm{\x}^2$
A penalty term is part of the objective, not a constraint: every $\x$ is still allowed.
a linear program
The objective has squares in it: it's quadratic, not linear.

The diet problem, and how the field began

In 1945 George Stigler asked for the cheapest diet that meets every nutritional need; it became one of the first problems solved by what we now call linear programming.

It shows, with a picture you can draw by hand, why constraints change everything: the answer sits on a corner, and the slope there is not zero.

Shopping on a budget: cheaper food is always tempting, but the rules (enough calories, enough protein) push you to a particular mix where both rules just barely hold.

A two-food version you can solve by eye

Using Stigler-style prices in dollars: oats cost $0.30 per serving and give 200 kcal and 10 g protein; milk costs $0.25 and gives 100 kcal and 10 g protein. You need at least 600 kcal and 40 g protein a day. Let $x_1$ be servings of oats and $x_2$ servings of milk.

$$\min_{\x}\ 0.30x_1+0.25x_2\quad\text{s.t.}\quad 2x_1+x_2\ge6,\quad x_1+x_2\ge4,\quad x_1\ge0,\ x_2\ge0.$$

(Calories in hundreds: $200x_1+100x_2\ge600$ becomes $2x_1+x_2\ge6$; protein: $10x_1+10x_2\ge40$ becomes $x_1+x_2\ge4$.) Both the objective and the constraints are linear: this is a linear program (LP).

Each constraint keeps one side of a straight line, so the feasible set is a region with straight edges, a polygon. The cost's level sets $0.30x_1+0.25x_2=C$ are parallel lines. Sliding that line towards lower cost until it's about to leave the feasible region, it last touches a corner (vertex). That's a general fact about linear programs: if an optimum exists, one is at a vertex.

Try it

The shaded region is every diet that meets both requirements. Drag the cost slider to slide the dashed iso-cost line; the cheapest diet is where it last touches the region. Then change the price of milk and watch the optimal corner jump.

Solve the two-food diet problem by checking the corners.

  1. Corners come from pairs of boundary lines meeting: $x_1=0$ with $2x_1+x_2=6$ gives $(0,6)$; the two nutrient lines $2x_1+x_2=6$, $x_1+x_2=4$ give $(2,2)$; $x_2=0$ with $x_1+x_2=4$ gives $(4,0)$.

    Subtracting the two nutrient equations gives $x_1=2$, then $x_2=2$. Other intersections, like $(3,0)$, violate a constraint, so they aren't corners of the feasible region.

  2. Costs: $(0,6)\to$1.50$, $(2,2)\to$1.10$, $(4,0)\to$1.20$.

    Evaluate the objective at each feasible corner.

  3. The cheapest is $\x^\star=(2,2)$: two servings of each, $1.10 a day. Both nutrient constraints hold with equality there (they're active).

    At the optimum, you're "just barely" meeting both requirements; any slack would be money wasted.

  4. Notice $\grad f=(0.30,0.25)^\top\ne\0$ everywhere. The minimizer is not where the slope vanishes.

    "Gradient equals zero" is the rule for unconstrained interior points. With constraints you need a different rule, which writes $\grad f$ as a nonnegative combination of the active constraints' gradients: here $(0.30,0.25)=0.05\,(2,1)+0.20\,(1,1)$. That rule (the KKT conditions) is beyond the lectures covered so far.

A short history (what the lecture covered)

WhenWhat happenedWhere you'll meet it
1600s–1700sFermat: at a smooth extremum the tangent is flat. Euler and Lagrange: calculus of variations; Lagrange multipliers (1797).Part 3 (first-order condition)
1847Cauchy proposes moving along $-\grad f$ to solve equations from astronomy: gradient descent.Parts 5–7
1939–1947Kantorovich (production planning) and Dantzig (air-force logistics) invent linear programming and the simplex method.Not covered here (Kantorovich's inequality in Part 5 is a different result of his)
1945Stigler's diet problem: 77 foods, 9 nutrients, solved approximately by hand.This chapter
1951Kuhn–Tucker (and Karush, 1939) conditions for constrained optima.Beyond Lectures 1–14
1959–1970Davidon, Fletcher–Powell, Broyden and others: quasi-Newton methods.Part 10
1984Karmarkar's interior-point method: fast, provably polynomial LP.—
1993Rockafellar: the great watershed is convexity vs non-convexity, not linearity vs non-linearity.Part 4
2010s–Machine learning makes cheap first-order methods (gradient descent and relatives) dominant.Parts 5–7

What is the cost of the diet $(x_1,x_2)=(2,2)$?

$0.30x_1+0.25x_2$.

$0.60+0.50=$1.10$.

Suppose milk drops to $0.10 per serving (oats stay at $0.30). Which corner is now cheapest? Enter $(x_1,x_2)$.

Recompute the cost at the three corners $(0,6)$, $(2,2)$, $(4,0)$ with the new price.

Costs: $(0,6)\to0.60$, $(2,2)\to0.80$, $(4,0)\to1.20$. The cheapest is now all milk: $(0,6)$. Changing prices tilts the iso-cost lines, so a different corner is touched last.

Is $(1,3)$ a feasible diet?

Check $2x_1+x_2\ge6$ and $x_1+x_2\ge4$ (and nonnegativity).

$2+3=5<6$: the calorie constraint fails. So it is not feasible.

  • Draw the feasible region when there are only two variables
  • For a linear program, check the corners (vertices)
  • Identify which constraints are active at the answer
  • Setting $\grad f=\0$ for a linear objective (it never is, unless $f$ is constant)
  • Including intersection points that violate some other constraint
  • Forgetting nonnegativity constraints like $x_i\ge0$
  1. The diet problem is a linear program: linear cost, linear constraints, polygon-shaped feasible set.
  2. For linear programs an optimum (if one exists) is at a vertex; at the optimum the gradient is not zero.
  3. Gradient descent (Cauchy, 1847) is one of the oldest algorithms in the field and is at the heart of modern machine learning.

Why can't we solve the diet problem by setting $\grad f(\x)=\0$?

Because the problem has no solution
It does: $(2,2)$ at $1.10.
Because $\grad f=(0.30,0.25)^\top$ is never zero; the minimum sits on the boundary, held there by constraints
"Gradient zero" only applies at interior points. Here the constraints stop you from going cheaper.
Because the objective isn't differentiable
A linear function is as smooth as can be.

Who first proposed gradient descent, and when?

Dantzig, 1947
Dantzig gave us the simplex method for linear programs.
Cauchy, 1847
To solve equations in astronomy, by following the negative gradient.
Rockafellar, 1993
Rockafellar's 1993 remark was about convexity being the real watershed.

Rockafellar's "great watershed" in optimization is between…

linear and nonlinear problems
That's exactly the older view he argued against.
convex and non-convex problems
Convex problems are tractable (local = global); non-convex ones generally aren't. Part 4.
constrained and unconstrained problems
Important, but not his point.

Classify the problem before you choose a method

No single algorithm works on every problem, so the first job is always to classify yours: convex or not, smooth or not, constrained or not.

A classic exam question hands you a function and asks which assumption fails and what goes wrong. This chapter rehearses exactly that.

A doctor diagnoses before prescribing: the same medicine that cures one illness can make another worse.

Problems are classified along four independent axes. Each one decides something concrete about which methods can work:

AxisClassesWhat it decides
Convexityconvex vs non-convexWhether a local answer is automatically global (Part 4). General non-convex problems are NP-hard.
Smoothness$C^2$, $C^1$, non-smoothNewton's method needs second derivatives; gradient descent needs first derivatives; kinks need special (subgradient) methods.
Constraintsnone, simple bounds, generalPlain line-search methods (this course) vs methods that handle constraints.
Formquadratic, linear, general nonlinearQuadratics: solve a linear system, or use conjugate gradients (Part 8). Linear: simplex. General: iterate.

To see why this matters, take the simplest algorithm, gradient descent with a fixed step $\alpha>0$: $$x_{k+1}=x_k-\alpha f'(x_k),$$ and run it blindly on three one-variable functions, each breaking a different assumption.

Gradient descent with $\alpha=0.8$ starts at $x_0=0.5$ on $f(x)=|x|$, whose minimizer is $x^\star=0$. What happens?

It converges to 0 in a few steps
That's what we'd hope, but look at the slope: it's $\pm1$ everywhere except at 0, and never shrinks.
It slowly creeps towards 0
Creeping would need the steps to get smaller, but $\alpha f'(x)=\pm0.8$ always.
It bounces between 0.5 and −0.3 forever
$0.5-0.8=-0.3$, then $-0.3+0.8=0.5$. A cycle.
Try it

Pick a function, set the step size and starting point, then press "Step" repeatedly (or "Run"). Each function breaks gradient descent in its own way. Watch the iterates and the table of values.

Three failures, three missing assumptions

  1. $f_1(x)=2x$ (linear, unconstrained). $f_1'\equiv2$, so $x_k=x_0-2k\alpha\to-\infty$ and $f_1(x_k)\to-\infty$. Nothing is wrong with the algorithm: the problem has no minimizer. It's unbounded below. Missing assumption: a minimizer exists.
  2. $f_2(x)=|x|$ (convex but not smooth). The slope is $\pm1$ right up to the kink at 0, so the step $\alpha f'(x)$ never shrinks and the iterates cycle (with $\alpha=0.8$ from $0.5$: $0.5\to-0.3\to0.5\to\cdots$). Missing assumption: differentiability (smoothness). Non-smooth problems need subgradient methods with step sizes that shrink to zero, which are beyond this course.
  3. $f_3(x)=x^4-2x^2$ (smooth but not convex). $f_3'(x)=4x(x^2-1)$ is zero at $-1,0,1$. Start exactly at $x_0=0$ and the gradient is 0, so the iterate never moves, even though 0 is a local maximum. Start at $x_0=0.01$ with a small step (say $\alpha=0.1$) and it rolls to $+1$; at $-0.01$, to $-1$. Missing assumption: convexity. Gradient descent can't tell a maximum, a saddle or a minimum apart by the gradient alone.

Run two steps of $x_{k+1}=x_k-\alpha f'(x_k)$ on $f(x)=x^4-2x^2$ from $x_0=0.5$ with $\alpha=0.1$.

  1. $f'(x)=4x^3-4x$. At $x_0=0.5$: $f'=4(0.125)-2=-1.5$.

    Negative slope means the function decreases to the right, so the step should move right.

  2. $x_1=0.5-0.1(-1.5)=0.65$.

    Subtracting $\alpha$ times a negative number moves us to the right, as expected.

  3. $f'(0.65)=4(0.274625)-2.6=1.0985-2.6=-1.5015$, so $x_2=0.65+0.15015=0.80015$.

    Still moving right, towards the minimizer at 1. As $x\to1$ the slope shrinks and so do the steps.

Gradient descent on $f(x)=|x|$ with $\alpha=0.8$, starting from $x_0=0.5$. What is $x_1$?

For $x>0$, $f'(x)=1$.

$x_1=0.5-0.8(1)=-0.3$. Then $x_2=-0.3-0.8(-1)=0.5$: back where we started.

Repeat the worked example yourself: $f(x)=x^4-2x^2$, $\alpha=0.1$, $x_0=0.5$. What is $x_2$?

$x_1=0.65$. Now compute $f'(0.65)=4(0.65)^3-4(0.65)$.

$f'(0.65)=-1.5015$, so $x_2=0.65+0.15015=0.80015$.

Which class does $f(x)=e^x+e^{-x}$ belong to?

Compute $f''(x)$. Is it always positive?

$f''(x)=e^x+e^{-x}>0$ everywhere and $f$ is infinitely differentiable: smooth and (strictly) convex. Its minimizer is $x=0$.

  • Before running anything, ask: does a minimizer exist? Is $f$ differentiable? Is it convex?
  • Match the method to the class: smooth → gradient methods; $C^2$ → Newton is possible; quadratic → conjugate gradients
  • Treat "the algorithm stopped" and "I found the minimum" as different claims
  • Running gradient descent on a function with kinks and trusting the output
  • Concluding a stationary point is a minimum without checking curvature or convexity
  • Blaming the step size when the real problem is that $\inf f=-\infty$
  1. Classify along four axes (convexity, smoothness, constraints, form) before choosing a method.
  2. Gradient descent fails on unbounded problems (runs off), on kinks (cycles), and can stall at non-minima of non-convex functions.
  3. Every convergence theorem later in the course comes with hypotheses; exam questions test whether you can spot which one fails.

Gradient descent with fixed step is started at $x_0=0$ on $f(x)=x^4-2x^2$. It…

moves to $x=1$
To move, it needs a nonzero gradient. What is $f'(0)$?
never moves, because $f'(0)=0$, though 0 is a local maximum
The update is $x_1=0-\alpha\cdot0=0$. The gradient alone can't tell a maximum from a minimum.
diverges to $-\infty$
$f$ is bounded below by $-1$; nothing runs away here.

Why does gradient descent cycle on $|x|$ instead of converging?

The step size is too small
A smaller fixed step would still cycle, just with a smaller swing: the slope never shrinks.
The slope stays $\pm1$ right up to the kink, so the steps never shrink
On smooth functions the slope fades near a minimum and the steps shrink with it. A kink breaks that.
$|x|$ is not convex
$|x|$ is convex. Its problem is that it's not differentiable at 0.

For a convex quadratic $\frac12\x^\top A\x-\b^\top\x$ with $A\succ0$, which classification is right?

Smooth, convex, unconstrained, quadratic
All the good boxes ticked: a unique global minimizer at $A^{-1}\b$, and every method in the course works on it.
Smooth, non-convex, unconstrained, quadratic
With $A\succ0$, the Hessian is positive definite: that's convex (Part 4).
Linear program
It has a quadratic term $\frac12\x^\top A\x$.

$x^\star$ is a local minimizer of $f$ on $\R$. Which must be true?

$f(x^\star)\le f(x)$ for all $x\in\R$
That's the definition of global.
There's an $\varepsilon>0$ with $f(x^\star)\le f(x)$ whenever $|x-x^\star|<\varepsilon$
"Local" means "best within some small ball".
$f''(x^\star)>0$
Not necessarily: $x^4$ has a minimizer at 0 with $f''(0)=0$. Part 3 covers this.

In 1-D, gradient descent moves $x$ in the direction of…

$-f'(x)$, i.e. right when the slope is negative
$x_{k+1}=x_k-\alpha f'(x_k)$: a negative slope pushes $x$ to the right, which is downhill.
$f'(x)$, i.e. right when the slope is positive
That's uphill. Descent goes against the slope.
always towards 0
It goes downhill, which needn't be towards 0.

In a linear program with a bounded, nonempty feasible polygon, an optimal solution…

is where the gradient of the cost is zero
A nonzero linear cost has a gradient that's never zero.
can be found at a vertex of the polygon
Slide the iso-cost line; it last touches the region at a corner (or along an edge, which includes corners).
is always in the interior
At an interior point you could always move to reduce a nonzero linear cost.

The ridge-regression matrix $A^\top A+\lambda I$ with $\lambda>0$ is positive definite because…

$\v^\top(A^\top A+\lambda I)\v=\norm{A\v}^2+\lambda\norm{\v}^2>0$ for $\v\ne\0$
The first term is $\ge0$ and the second is $>0$ for any nonzero $\v$.
all its entries are positive
Positive entries don't imply positive definiteness, and $A^\top A$ can have negative entries anyway.
it is symmetric
Symmetric matrices can be indefinite.

Lectures 2 and 3 (11 and 13 August), taken from Rudin's Principles of Mathematical Analysis: the supremum and infimum, distance and balls, open and closed sets, limit points and closure, convergent sequences, continuity, compactness, and the payoff, the Weierstrass extreme value theorem. We finish with coercivity, the tool that guarantees a minimizer on all of $\R^n$. The diary says the proofs are "good to know" but not needed, so the proofs of Rudin's results sit in optional "Go deeper" boxes. The definitions, the theorem statements, the examples and the counterexamples are all examinable, and they're all in the main text.

You need: Part 0a (sets, quantifiers $\forall,\exists$, min vs inf, norms, level sets), Part 0b (eigenvalues and positive definiteness, for the last chapter) and Part 1 (local and global minimizers, the ball $B(\x^\star,\varepsilon)$).

Supremum and infimum: the best bound, even when nobody reaches it

The infimum of a set of numbers is its greatest lower bound. It always exists for a nonempty set that's bounded below, even when the set has no smallest element.

The optimal value $p^\star=\inf_{\x\in\mathcal F}f(\x)$ is defined this way. A minimizer exists exactly when this infimum is actually reached, so "inf versus min" is the first question to settle in any optimization problem.

A high-jump bar you can get arbitrarily close to clearing but never quite clear: the bar height is well defined (it's the supremum of your jumps), even though no single jump equals it.

Try to minimize $f(x)=x$ over the open interval $(0,1)$. Every candidate $x$ loses to $x/2$, which is smaller and still in $(0,1)$. So there is no smallest value. Yet the number 0 still plays a special role as a bound: no value is below 0, and the values get as close to 0 as you like. We need a word for that number. It's the infimum.

Where this happens for real

Logistic regression on separable data. A classifier with weight $w$ on a single correctly labelled example has loss $\ell(w)=\log(1+e^{-w})$. As $w$ grows, the loss shrinks towards 0, but $\ell(w)>0$ for every $w$. So $\inf\ell=0$ and no minimizer exists. Training software running gradient descent here pushes $w$ off towards infinity forever. This is a real, practical failure, and it's the reason regularization (adding $\frac\lambda2w^2$) is used: the regularized loss has a minimizer, as the last chapter shows.

Bounds

Let $E$ be a set of real numbers. A number $\beta$ is an upper bound of $E$ if $x\le\beta$ for every $x\in E$. If $E$ has an upper bound, $E$ is bounded above. Lower bounds and "bounded below" are defined the same way with $\ge$. A set can have many upper bounds: for $E=(0,1)$, the numbers $1$, $2$ and $100$ are all upper bounds. The interesting one is the smallest.

Let $E\subseteq\R$ be bounded above. A number $\alpha$ is the least upper bound or supremum of $E$, written $\alpha=\sup E$, if

  1. $\alpha$ is an upper bound of $E$, and
  2. if $\gamma\lt\alpha$, then $\gamma$ is not an upper bound of $E$. Equivalently: for every $\varepsilon>0$ there is some $x\in E$ with $x>\alpha-\varepsilon$.

The greatest lower bound or infimum, $\inf E$, is defined in the same way: $\alpha=\inf E$ means $\alpha$ is a lower bound and no $\beta>\alpha$ is a lower bound (for every $\varepsilon>0$ some $x\in E$ has $x\lt\alpha+\varepsilon$).

Read condition (ii) as a challenge game. Someone proposes a number $\gamma$ slightly below $\alpha$. You must answer with an element of $E$ that beats $\gamma$. If you can always answer, however close $\gamma$ is to $\alpha$, then $\alpha$ is the least upper bound. Condition (ii) is the one students forget when asked to prove that a number is a supremum.

Max versus sup. A maximum of $E$ is an upper bound that belongs to $E$. If $E$ has a maximum, then $\max E=\sup E$. The supremum may or may not belong to $E$: the maximum exists exactly when $\sup E\in E$. The same holds for minimum and infimum. In optimization language: a minimizer exists exactly when the infimum is attained.

Set $E$supinf
$(0,1]$1, attained0, not attained
$\{1-\frac1n\}$: $0,\frac12,\frac23,\dots$1, not attained0, attained ($n=1$)
$\{\frac1n\}$ (Rudin 1.9c)1, attained ($n=1$)0, not attained
$\{\frac{(-1)^n}n\}$: $-1,\frac12,-\frac13,\dots$$\frac12$, attained ($n=2$)$-1$, attained ($n=1$)
$\{r\in\Q:r\lt0\}$ (Rudin 1.9b)0, not attainednone
$\{r\in\Q:r\le0\}$0, attainednone

The terms $\frac{(-1)^n}{n}$ get closer and closer to 0. Is the supremum of $\{\frac{(-1)^n}n:n\in\N\}$ attained?

No: the sup is 0, which is approached but never reached
0 is the limit of the terms, not their supremum. Is 0 even an upper bound? Look at $n=2$.
Yes: the sup is $\frac12$, reached at $n=2$
The positive terms are $\frac12,\frac14,\frac16,\dots$, and the largest of them is $\frac12$. The limit of a sequence and the sup of its terms are different things.
No: the set has no upper bound
Every term is at most $\frac12$ in absolute value after $n=1$, and $-1$ is negative. It's bounded.
Try it

Pick a set and slide the candidate $\gamma$. In upper-bound mode, the widget either confirms $\gamma$ is an upper bound or produces an element of $E$ that beats it (ringed, with its index $n$). Put $\gamma$ just below the sup of $\{1-\frac1n\}$, at 0.999, and see how large $n$ must be. Then do the same in lower-bound mode for $\inf\,(0,1]$.

Completeness: why we work in $\R$ and not in $\Q$

Does every bounded set have a supremum? In the rationals $\Q$, no. Take $E=\{p\in\Q:p>0,\ p^2\lt2\}$. It's bounded above by 2, but its least upper bound "should" be $\sqrt2$, which isn't rational (Part 0a's $\Q$ is fractions only). Whatever rational upper bound you propose, a slightly smaller rational upper bound exists. So inside $\Q$, $E$ has no supremum at all.

An ordered set $S$ has the least-upper-bound property if every nonempty $E\subseteq S$ that is bounded above has a supremum in $S$.

$\R$ has the least-upper-bound property: every nonempty set of reals that is bounded above has a real supremum (and every nonempty set bounded below has a real infimum).

This is the foundation of the whole existence theory. For a nonempty feasible set $\mathcal F$, the optimal value $p^\star=\inf_{\x\in\mathcal F}f(\x)$ always makes sense: it's a real number if $f$ is bounded below on $\mathcal F$, and we write $p^\star=-\infty$ if it isn't. By convention $\inf_{\emptyset}f=+\infty$ for an infeasible problem. (Allowing $\pm\infty$ as values is called the extended real line.) So the value always exists. The minimizer is what may fail to exist, and the rest of this part is about when it's guaranteed.

One consequence of completeness gets used in nearly every proof with an $\varepsilon$ in it:

For every real $\varepsilon>0$ there is a natural number $n$ with $\frac1n\lt\varepsilon$ (equivalently $n\varepsilon>1$). More generally, for $x>0$ and any real $y$ there is $n\in\N$ with $nx>y$.

Prove from the definition that $\sup\{1-\frac1n:n\in\N\}=1$, and show that this supremum is not attained.

  1. (i) 1 is an upper bound. For every $n$, $\frac1n>0$, so $1-\frac1n\lt1$.

    Condition (i) of the definition: check every element of the set.

  2. (ii) Nothing smaller is an upper bound. Let $\gamma\lt1$, so $1-\gamma>0$. By the Archimedean property there is $n$ with $\frac1n\lt1-\gamma$. Then $1-\frac1n>\gamma$: an element of $E$ beats $\gamma$.

    Condition (ii): for any challenger $\gamma$ below 1, we exhibit a specific element above it. For $\gamma=0.999$, $1-\gamma=0.001$ and $n=1001$ works.

  3. By (i) and (ii), $\sup E=1$.

    Both conditions together are the definition.

  4. Not attained: $1-\frac1n=1$ would need $\frac1n=0$, impossible. So $1\notin E$ and $E$ has no maximum.

    The supremum exists (completeness) but isn't a member of the set: exactly the "inf without a minimizer" situation, flipped.

Go deeper: proofs of the Archimedean property and of $\sqrt2\notin\Q$

Archimedean property. Suppose $nx\le y$ for every $n$. Then $A=\{nx:n\in\N\}$ is nonempty and bounded above by $y$, so $\alpha=\sup A$ exists by completeness. Since $x>0$, $\alpha-x\lt\alpha$ is not an upper bound, so $\alpha-x\lt mx$ for some $m$. Then $\alpha\lt(m+1)x\in A$, contradicting that $\alpha$ is an upper bound.

No rational square root of 2. Suppose $(a/b)^2=2$ with $a/b$ in lowest terms. Then $a^2=2b^2$ is even, so $a$ is even, $a=2c$. Then $4c^2=2b^2$, so $b^2=2c^2$ and $b$ is even too, contradicting lowest terms.

Uniqueness and the sup–inf link. A set has at most one supremum: if $\alpha\lt\alpha'$ were both suprema, $\alpha$ would be an upper bound smaller than $\alpha'$, contradicting (ii) for $\alpha'$. Also $\inf E=-\sup(-E)$, where $-E=\{-x:x\in E\}$. That's why minimizing $f$ and maximizing $-f$ are the same problem.

How $\R$ is actually built from $\Q$ (Dedekind cuts, Rudin's appendix to Chapter 1) is not examined.

Let $E=\{(-1)^n\,(1-\frac1n):n\in\N\}$. Find $\sup E$ and $\inf E$. (Then decide for yourself whether either is attained.)

List the first few terms: $n=1,2,3,4,5$ give $0,\ \frac12,\ -\frac23,\ \frac34,\ -\frac45$. Even $n$ give positive terms creeping up towards what?

Even $n$: $1-\frac1n\in\{\frac12,\frac34,\frac56,\dots\}$, all below 1 and getting within any $\varepsilon$ of 1. Odd $n$: $-(1-\frac1n)\in\{0,-\frac23,-\frac45,\dots\}$, all above $-1$ and getting within any $\varepsilon$ of $-1$. So $\sup E=1$ and $\inf E=-1$. Neither is attained, because $1-\frac1n\ne1$ for every $n$. So $E$ has neither a maximum nor a minimum, although it's bounded.

For $E=\{1-\frac1n\}$ with $\sup E=1$, the challenger is $\gamma=1-0.003$. What is the smallest $n$ whose element $1-\frac1n$ beats $\gamma$?

$1-\frac1n>1-0.003$ is the same as $\frac1n\lt0.003$, i.e. $n>\frac1{0.003}$.

$\frac1n\lt0.003\iff n>333.33\ldots$, so the smallest such $n$ is $334$. (Check: $\frac1{333}\approx0.003003>0.003$ fails; $\frac1{334}\approx0.002994\lt0.003$ works.)

Find $\inf\{x+\frac1x : x>0\}$.

$x+\frac1x-2=\frac{(x-1)^2}{x}$. What sign does that have for $x>0$?

For $x>0$, $x+\frac1x-2=\frac{(x-1)^2}x\ge0$, so 2 is a lower bound. It's attained at $x=1$, so no larger number can be a lower bound. Hence $\inf=\min=2$.

$E\subseteq\R$ is nonempty and bounded above, and $\sup E\notin E$. Which statement must be true?

What is $\sup E$ for a finite set like $\{2,5,3\}$? Does it belong to the set?

A finite nonempty set has a largest element, and that element is its supremum, so the sup is in the set. Since here $\sup E\notin E$, $E$ must be infinite. The others can fail: $E=\{1-\frac1n\}$ isn't open, has a minimum (0), and its sup (1) is rational.

  • To prove $\alpha=\sup E$, check both parts: $\alpha$ is an upper bound, and every $\gamma\lt\alpha$ is beaten by some element
  • Ask "is the inf attained?" before writing "min"
  • Write $p^\star=-\infty$ for a problem unbounded below, and $+\infty$ for an infeasible one
  • List the first few terms of a sequence before guessing its sup and inf
  • Confusing the limit of a sequence with the sup of its terms ($\{(-1)^n/n\}$ has limit 0 but sup $\frac12$)
  • Proving only that $\alpha$ is an upper bound and stopping
  • Saying "the minimum is 0" for $f(x)=e^{-x}$ on $[0,\infty)$: the infimum is 0, and there's no minimum
  1. $\sup E$ is the least upper bound: an upper bound that every smaller number fails to be. $\inf E$ is the mirror image.
  2. $\R$ is complete: every nonempty bounded-above set has a supremum. So $p^\star=\inf_{\mathcal F}f$ always exists (possibly $-\infty$).
  3. A max/min exists exactly when the sup/inf is attained, i.e. belongs to the set. Optimization theory is about when that's guaranteed.

$\alpha$ is an upper bound of $E$. To conclude $\alpha=\sup E$, you must also show:

$\alpha\in E$
That would make $\alpha$ a maximum. A supremum needn't belong to the set: think of $(0,1)$.
for every $\varepsilon>0$ some $x\in E$ has $x>\alpha-\varepsilon$
That's condition (ii): nothing smaller than $\alpha$ is still an upper bound.
$\alpha$ is larger than every element of $E$
That's just (a strict version of) being an upper bound, which you already have. What makes it the least one?

Why is $\{p\in\Q: p>0,\ p^2\lt2\}$ the standard example in Rudin's Chapter 1?

It is not bounded above
2 is an upper bound.
It has a maximum in $\Q$
For any rational $p$ with $p^2\lt2$ there's a slightly larger rational with the same property.
It is bounded above but has no least upper bound in $\Q$, so $\Q$ lacks the least-upper-bound property
Its sup "should" be $\sqrt2$, which isn't rational. $\R$ fixes exactly this gap.

For $\min_{x\in\R}\ e^{x}$, which is correct?

$p^\star=-\infty$
$e^x>0$ for every $x$, so 0 is a lower bound and $p^\star$ is finite.
$p^\star=0$, and no minimizer exists
$e^x\to0$ as $x\to-\infty$ but never equals 0: the inf isn't attained.
$p^\star=0$, attained as $x\to-\infty$
"Attained" means some actual $x$ in the domain gives $f(x)=p^\star$. $-\infty$ isn't a real number.

A set $E$ is nonempty and bounded below, and $\inf E\in E$. Then:

$E$ has a minimum, equal to $\inf E$
A lower bound that belongs to the set is the smallest element.
$E$ must be closed
Try $E=\{0\}\cup(1,2)$: its inf 0 is in $E$, but $E$ is not closed (2 is a missing limit point).
$E$ must be finite
$[0,1]$ is infinite and contains its inf.

Distance, balls, and open and closed sets

A set is closed if it contains every point it can be approached from, and open if every point in it has some room around it.

Minimizers hide on edges. A feasible set that includes its edges ($\le$ constraints) can hold the minimizer there; a set that leaves them out ($\lt$ constraints) can let it slip away, as $(0,1)$ did in Chapter 1.

A field with a fence: if the fence line belongs to the field (closed), you can stand on it. If it doesn't (open), you can get as close as you like, but every spot you stand on still has a little grass between you and the fence.

Distance and metric spaces

Everything in this chapter is built from one idea: the distance between two points. In $\R^n$ it's the Euclidean distance $d(\x,\y)=\norm{\x-\y}=\sqrt{\sum_i(x_i-y_i)^2}$. Rudin's approach is to list the three properties of distance that every argument actually uses. Any set with a distance function obeying them is a metric space, and everything below works there.

A set $X$ (whose elements are called points) is a metric space if there is a function $d:X\times X\to\R$, the distance or metric, such that for all $p,q,r\in X$:

  1. $d(p,q)>0$ if $p\ne q$, and $d(p,p)=0$;
  2. $d(p,q)=d(q,p)$ (symmetry);
  3. $d(p,q)\le d(p,r)+d(r,q)$ (the triangle inequality: a detour through $r$ is never shorter).

For $\x,\y,\z\in\R^k$ and $\lambda\in\R$: (a) $\norm\x\ge0$, with equality only for $\x=\0$; (b) $\norm{\lambda\x}=|\lambda|\,\norm\x$; (c) $\norm{\x+\y}\le\norm\x+\norm\y$; (d) $\norm{\x-\z}\le\norm{\x-\y}+\norm{\y-\z}$.

Consequently (Rudin's Remark 1.38 and Example 2.16), $\R^k$ with $d(\x,\y)=\norm{\x-\y}$ is a metric space. Every subset of a metric space, such as a feasible set $\mathcal F\subseteq\R^n$, is a metric space with the same $d$.

Prove part (c), $\norm{\x+\y}\le\norm\x+\norm\y$, and deduce (d). (The book's exam tip says this two-line argument is the expected answer to "justify the triangle inequality".)

  1. Expand: $\norm{\x+\y}^2=(\x+\y)^\top(\x+\y)=\norm\x^2+2\,\x^\top\y+\norm\y^2$.

    The squared norm is the dot product of a vector with itself, and the dot product distributes like ordinary multiplication.

  2. By Cauchy–Schwarz, $\x^\top\y\le|\x^\top\y|\le\norm\x\norm\y$. So $\norm{\x+\y}^2\le\norm\x^2+2\norm\x\norm\y+\norm\y^2=(\norm\x+\norm\y)^2$.

    Cauchy–Schwarz (Part 0a) is the only real input. The right side is now a perfect square.

  3. Both sides are nonnegative, so taking square roots keeps the inequality: $\norm{\x+\y}\le\norm\x+\norm\y$.

    For $a,b\ge0$, $a^2\le b^2$ implies $a\le b$.

  4. For (d), apply (c) to the vectors $\x-\y$ and $\y-\z$, whose sum is $\x-\z$.

    This turns the "vector" inequality (c) into the "distance" inequality (d), which is axiom (c) of a metric space.

The 1-norm and $\infty$-norm of Part 0a also satisfy the triangle inequality, so $d_1(\x,\y)=\norm{\x-\y}_1$ and $d_\infty(\x,\y)=\norm{\x-\y}_\infty$ are metrics on $\R^n$ too. Their "balls" are diamonds and squares instead of discs. In $\R^n$ all three give the same open sets, closed sets and convergent sequences, so for this course the choice doesn't matter.

The vocabulary of Rudin's Definition 2.18

From now on, $X$ is a metric space (think $\R^2$) and $E\subseteq X$. The basic object is the neighbourhood of $p$ with radius $r>0$: $$N_r(p)=\{q\in X: d(p,q)\lt r\},$$ the open ball of Part 1, $B(p,r)$. In $\R$ it's the interval $(p-r,p+r)$; in $\R^2$ it's a disc without its rim.

  1. A neighbourhood of $p$ is a set $N_r(p)$ for some radius $r>0$.
  2. $p$ is a limit point of $E$ if every neighbourhood of $p$ contains a point $q\ne p$ with $q\in E$. ($p$ itself may or may not be in $E$.)
  3. If $p\in E$ and $p$ is not a limit point of $E$, then $p$ is an isolated point of $E$.
  4. $E$ is closed if every limit point of $E$ is a point of $E$.
  5. $p$ is an interior point of $E$ if there is a neighbourhood $N$ of $p$ with $N\subseteq E$.
  6. $E$ is open if every point of $E$ is an interior point of $E$.
  7. The complement of $E$ is $E^c=\{p\in X: p\notin E\}$.
  8. $E$ is perfect if $E$ is closed and every point of $E$ is a limit point of $E$.
  9. $E$ is bounded if there are a real $M$ and a point $q\in X$ with $d(p,q)\lt M$ for all $p\in E$.
  10. $E$ is dense in $X$ if every point of $X$ is a limit point of $E$, or a point of $E$, or both.
Each definition in plain words
limit point
You can find points of $E$ (other than $p$) arbitrarily close to $p$. Shrinking the ball never empties it.
isolated point
A point of $E$ standing alone: some small ball around it contains no other point of $E$.
interior point
A point with breathing room: some ball around it lies entirely inside $E$.
closed
$E$ contains everything it can be approached from.
open
Every point of $E$ has breathing room.
perfect
Closed, with no isolated points ($[0,1]$ is perfect; $[0,1]\cup\{2\}$ is not).
bounded
$E$ fits inside some ball.
dense
$E$ gets arbitrarily close to every point of $X$: $\Q$ is dense in $\R$.
boundary point
(Not in Rudin's list, but handy.) Every ball around $p$ meets both $E$ and $E^c$. The unit circle is the boundary of the unit disc.
Careful: "open" and "closed" are not opposites. $[0,1)$ is neither (the point 1 is a limit point not in the set, and 0 has no breathing room). $\R^n$ and $\emptyset$ are both. Most sets are neither. What is true is the complement rule: $E$ is open if and only if $E^c$ is closed ([Rudin] Thm. 2.23).
Try it

Choose a set, click to place a point $p$, and shrink the ball radius $r$ towards 0. For each $r$ the readout answers two questions: does the ball fit inside $E$ (interior test) and does it contain a point of $E$ other than $p$ (limit-point test)? Decide what $p$ is before revealing the verdict. Test the centre of the punctured disc, a point on the circle of the open disc, a dot of the sequence, the origin next to the sequence, and a point on the segment.

Two theorems from Lecture 2 make the vocabulary consistent:

Every neighbourhood $N_r(p)$ is an open set.

If $p$ is a limit point of $E$, then every neighbourhood of $p$ contains infinitely many points of $E$.

Corollary. A finite set has no limit points, so every finite set is closed.

Theorem 2.19 says the name "open ball" is honest. Theorem 2.20 says a limit point is not an accident of one nearby point: shrink the ball and new points of $E$ keep appearing inside. That's why the limit point of $\{\frac1n\}$ is 0 (infinitely many terms crowd near it) and why a single point like $\frac13$ is isolated.

Go deeper: proofs of Theorems 2.19 and 2.20

2.19. Let $q\in N_r(p)$, so $d(p,q)=r-h$ for some $h>0$. If $d(q,s)\lt h$, then by the triangle inequality $d(p,s)\le d(p,q)+d(q,s)\lt(r-h)+h=r$, so $s\in N_r(p)$. Thus $N_h(q)\subseteq N_r(p)$: $q$ is an interior point. The leftover distance $h=r-d(p,q)$ is the breathing room.

2.20. Suppose some neighbourhood $N$ of $p$ contains only finitely many points of $E$. Let $q_1,\dots,q_n$ be those points other than $p$, and let $r=\min_m d(p,q_m)>0$ (a minimum of finitely many positive numbers is positive). Then $N_r(p)$ contains no point of $E$ other than $p$, so $p$ is not a limit point.

Infinite intersections ([Rudin] Ex. 2.25). Any union of open sets is open, and a finite intersection of open sets is open. But $\bigcap_{n\ge1}(-\frac1n,\frac1n)=\{0\}$, which is not open. Dually, any intersection of closed sets is closed, but an infinite union of closed sets need not be: $\bigcup_{n\ge1}[\frac1n,1]=(0,1]$.

Closure: adding the missing limit points

Let $E'$ be the set of all limit points of $E$. The closure of $E$ is $\bar E=E\cup E'$.

([Rudin] Thm. 2.27) $\bar E$ is closed; $E=\bar E$ exactly when $E$ is closed; and $\bar E$ is the smallest closed set containing $E$.

Examples: the closure of $(0,1)$ is $[0,1]$; of the open unit disc, the closed unit disc; of $\{\frac1n\}$, the set $\{\frac1n\}\cup\{0\}$; of $\Q$ in $\R$, all of $\R$ (that's what "dense" means).

Now the link to Chapter 1. The supremum of a set can always be approached from inside the set, so it's either in the set or a limit point of it:

Let $E\subseteq\R$ be nonempty and bounded above, and $y=\sup E$. Then $y\in\bar E$. Hence $y\in E$ if $E$ is closed. (The same holds for $\inf E$ when $E$ is bounded below.)

So a closed, bounded set of real numbers always has a maximum and a minimum. That's the seed of the Weierstrass theorem in Chapter 4: if we can show that the set of values $\{f(\x):\x\in K\}$ is closed and bounded, its infimum is attained, and that is a minimizer.

Go deeper: proof of Theorem 2.28

If $y\in E$ there is nothing to prove. Otherwise, take any $h>0$. Since $y-h\lt y=\sup E$, the number $y-h$ is not an upper bound, so some $x\in E$ has $y-h\lt x$, and $x\le y$ with $x\ne y$ because $y\notin E$. So every neighbourhood $(y-h,y+h)$ contains a point of $E$ other than $y$: $y$ is a limit point, hence in $\bar E$.

Let $E=\{\frac1n:n\in\N\}\subseteq\R$. Find its limit points and isolated points, decide whether it's open, closed, perfect or bounded, find $\bar E$, and check Theorem 2.28 for $\inf E$.

  1. 0 is a limit point. Given any $r>0$, the Archimedean property gives $n$ with $\frac1n\lt r$, so $N_r(0)=(-r,r)$ contains $\frac1n\ne0$.

    Definition (b): every neighbourhood of 0 must contain a point of $E$ other than 0. In fact it contains infinitely many (all $\frac1m$ with $m\ge n$), as Theorem 2.20 promises.

  2. Every point of $E$ is isolated. The nearest other point to $\frac1n$ is $\frac1{n+1}$, at distance $\frac1n-\frac1{n+1}=\frac1{n(n+1)}$. A ball of that radius around $\frac1n$ contains no other point of $E$.

    Definition (c). The ball can be chosen small enough to exclude all neighbours because there are none closer than $\frac1{n(n+1)}$.

  3. No other limit points. A point $p\lt0$ or $p>1$ has a ball missing $E$ altogether; a point $p\in(0,1]$ not in $E$ lies strictly between two consecutive terms and has a ball missing $E$.

    So $E'=\{0\}$.

  4. Not closed: the limit point 0 is not in $E$. Not open: no ball around $\frac12$ fits inside $E$ (it would contain non-members like 0.5001). Not perfect (not closed, and all its points are isolated). Bounded: $E\subseteq N_2(0)$.

    Each property is checked straight from its definition.

  5. $\bar E=E\cup\{0\}$, which is closed. $\inf E=0\in\bar E$ but $0\notin E$: Theorem 2.28 holds, and $E$ has no minimum.

    The infimum is a limit point outside the set, which is exactly how a minimum fails to exist.

The point $q=(0.3,0.4)$ lies in the open unit disc $N_1(\0)\subseteq\R^2$. What is the largest $h$ such that $N_h(q)\subseteq N_1(\0)$?

This is the breathing room in the proof of Theorem 2.19: $h=r-d(p,q)$. What is $\norm{(0.3,0.4)}$?

$\norm q=\sqrt{0.09+0.16}=0.5$, so $h=1-0.5=0.5$. Any larger radius pokes out past the circle along the direction of $q$, at the point $q+h\,q/\norm q$.

Which points are limit points of $E=\{(-1)^n+\frac1n : n\in\N\}\subseteq\R$?

Separate even and odd $n$. Even: $1+\frac1n$. Odd: $-1+\frac1n$. Where do the points crowd?

Even $n$ give $1+\frac12,1+\frac14,\dots$, crowding at 1; odd $n$ give $0,-1+\frac13,-1+\frac15,\dots$, crowding at $-1$. Every ball around 1 or $-1$ contains infinitely many of these points, so both are limit points. Neither is in $E$, so $E$ is not closed. (0 is in $E$, at $n=1$, but it's isolated.)

Is the half-disc $E=\{(x,y): x^2+y^2\le1,\ y>0\}$ open, closed, both or neither in $\R^2$?

Test the point $(0,1)$ (top of the arc) for openness, and the point $(0,0)$ for closedness.

$(0,1)\in E$, but every ball around it contains points with $x^2+y^2>1$, so it's not interior: $E$ is not open. $(0,0)\notin E$ (since $y=0$), but every ball around it contains points of $E$ like $(0,\frac r2)$: it's a limit point outside $E$, so $E$ is not closed. Neither.

Let $E=\{x\in\Q: 0\lt x\lt1\}$, a subset of $\R$. What is its closure $\bar E$?

Is an irrational number like $\frac{\sqrt2}{2}$ a limit point of $E$? (Use that $\Q$ is dense in $\R$.) What about 0 and 1?

Between any two reals there's a rational, so every real $p\in[0,1]$ has rationals of $(0,1)$ arbitrarily close to it, other than itself: every point of $[0,1]$ is a limit point. Points outside $[0,1]$ have a ball missing $(0,1)$ entirely. So $\bar E=[0,1]$. A set full of "holes" can still have a solid closure.

  • Test "closed" by hunting for a limit point outside the set, and "open" by hunting for a point with no breathing room
  • Remember that sets defined by $\le$ or $=$ of continuous functions are closed, and by strict $\lt$ are open (proved in the next chapter)
  • Use the closure to describe exactly what a set is missing
  • Thinking "not open" means "closed"
  • Forgetting that a limit point needn't belong to the set
  • Calling a point of $E$ a limit point just because it's in $E$ (isolated points aren't)
  • Judging interior points without saying which space you're in: $[-1,1]$ has interior points in $\R$, but as a segment in $\R^2$ it has none
  1. A metric is a distance obeying positivity, symmetry and the triangle inequality; $\norm{\x-\y}$ on $\R^n$ is one (Rudin 1.37).
  2. Closed = contains all its limit points; open = every point is interior. They are not opposites, but $E$ is open exactly when $E^c$ is closed.
  3. The closure $\bar E=E\cup E'$ adds the missing limit points, and $\sup E,\inf E\in\bar E$ (Rudin 2.28): on a closed set, the sup and inf are attained.

$p$ is a limit point of $E$. Which must be true?

$p\in E$
The centre of the punctured disc is a limit point but not a member.
Some neighbourhood of $p$ contains exactly one point of $E$ other than $p$
Theorem 2.20 says the opposite: every neighbourhood contains infinitely many.
Every neighbourhood of $p$ contains infinitely many points of $E$
Rudin's Theorem 2.20. In particular, finite sets have no limit points.

Which set is closed in $\R$?

$\{\frac1n : n\in\N\}$
0 is a limit point that's missing.
$\Z$, the integers
Every integer is isolated and $\Z$ has no limit points at all (any $p$ has a ball containing at most one integer), so it vacuously contains all of them. It's closed but unbounded.
$\Q$
Every real number is a limit point of $\Q$, and most aren't rational.

The feasible set $\{x\in\R : x>0\}$ is…

open, not closed
Each $x>0$ has the ball $(x/2,\,3x/2)$ inside the set; 0 is a limit point outside it. So $\min_{x>0}x$ has no solution.
closed, not open
Is 0 a limit point? Is it in the set?
both open and closed
In $\R$ only $\emptyset$ and $\R$ are both.

$E\subseteq\R$ is closed, nonempty and bounded above. What does Rudin's Theorem 2.28 give?

$E$ is compact
That needs bounded below as well (Chapter 4), and it's not what 2.28 says.
$\sup E\in E$, so $E$ has a maximum
$\sup E\in\bar E=E$ because $E$ is closed.
$\sup E$ is a limit point of $E$
Not necessarily: in $E=\{0\}\cup\{5\}$, the sup 5 is isolated. It's in $\bar E$ either way.

Sequences, continuity, and the intermediate value theorem

A sequence converges to $L$ if, for every tolerance you name, the terms eventually stay within that tolerance of $L$; a function is continuous if small enough input changes cause output changes within any tolerance you name.

Every algorithm in this course produces a sequence $\x_0,\x_1,\x_2,\dots$, and every convergence theorem is a statement in exactly this $\varepsilon$–$N$ language. Continuity is the hypothesis behind Weierstrass in the next chapter.

A thermostat: "eventually within half a degree, and staying there" is convergence. A dimmer switch is continuous (turn it a little, the light changes a little); a light switch is not.

A sequence $\{p_n\}$ in a metric space $X$ converges to $p\in X$ if for every $\varepsilon>0$ there is an integer $N$ such that $n\ge N$ implies $d(p_n,p)\lt\varepsilon$. We write $p_n\to p$ or $\lim_{n\to\infty}p_n=p$. If no such $p$ exists, the sequence diverges.

Read it as a contract. The sceptic names a tolerance $\varepsilon$; you must answer with an index $N$ after which every term is within $\varepsilon$ of $p$. The order of the quantifiers matters (Part 0a): $N$ is chosen after $\varepsilon$, so it may depend on it, and smaller tolerances usually need larger $N$. Limits are unique (two different limits $p\ne q$ would need terms within $\frac12d(p,q)$ of both, contradicting the triangle inequality).

Try it

Pick a sequence. The green band is $L\pm\varepsilon$ from index $N$ onwards. Shrink $\varepsilon$ and press the button to find the smallest $N$ that works. Then pick $a_n=(-1)^n$ and try every candidate $L$ you like with $\varepsilon=0.5$: no $N$ ever works, which is what "diverges" means. Finally try $n/(n+1)$ with the wrong limit $L=0.99$.

Prove that $a_n=\frac n{n+1}\to1$, and find the smallest valid $N$ for $\varepsilon=0.01$. Then show $b_n=(-1)^n$ diverges.

  1. $|a_n-1|=\left|\frac n{n+1}-1\right|=\frac1{n+1}$.

    Always start by simplifying the distance $|a_n-L|$ into something whose size you can read off.

  2. Given $\varepsilon>0$, choose $N$ with $N+1>\frac1\varepsilon$ (possible by the Archimedean property). For $n\ge N$: $\frac1{n+1}\le\frac1{N+1}\lt\varepsilon$.

    $\frac1{n+1}$ decreases in $n$, so if index $N$ is inside the band, every later index is too. This proves $a_n\to1$.

  3. For $\varepsilon=0.01$: $\frac1{n+1}\lt0.01\iff n+1>100\iff n\ge100$. The smallest $N$ is $100$.

    Check the edge: $n=99$ gives exactly $\frac1{100}=0.01$, which is not $\lt0.01$.

  4. For $b_n=(-1)^n$: suppose $b_n\to L$. With $\varepsilon=1$, eventually $|1-L|\lt1$ and $|-1-L|\lt1$, so $2=|1-(-1)|\le|1-L|+|L-(-1)|\lt2$. Contradiction.

    To prove divergence, show that one specific $\varepsilon$ defeats every candidate $L$. The triangle inequality does the work.

Sequences describe closed sets and limit points

The $\varepsilon$-ball definitions of Chapter 2 have sequence versions, and in practice these are what you use:

  • $p$ is a limit point of $E$ if and only if some sequence of points of $E\setminus\{p\}$ converges to $p$.
  • $E$ is closed if and only if whenever a sequence of points of $E$ converges, its limit is in $E$. ("You can't escape a closed set by taking limits.")

Example: the sequence $\frac1n\in(0,1]$ converges to $0\notin(0,1]$, so $(0,1]$ isn't closed. An algorithm whose iterates stay feasible and converge, on a closed feasible set, converges to a feasible point.

Continuity

Let $X,Y$ be metric spaces, $E\subseteq X$, $p\in E$ and $f:E\to Y$. Then $f$ is continuous at $p$ if for every $\varepsilon>0$ there is a $\delta>0$ such that $$d_Y(f(x),f(p))\lt\varepsilon\quad\text{for all }x\in E\text{ with }d_X(x,p)\lt\delta.$$ $f$ is continuous on $E$ if it's continuous at every point of $E$. Note that $f$ must be defined at $p$ to be continuous there.

$f$ is continuous at $p$ if and only if $f(x_n)\to f(p)$ for every sequence $x_n\to p$ in $E$. So to prove $f$ is discontinuous at $p$, exhibit one sequence $x_n\to p$ with $f(x_n)\not\to f(p)$.

Building blocks: constants, $x\mapsto x_i$, sums, products, quotients (where the denominator isn't 0), and compositions of continuous functions are continuous (Rudin Thms. 4.7, 4.9). So polynomials, norms, $e^x$, $\log$ on $(0,\infty)$, and every objective in this course are continuous.

Try it

The green horizontal band is $f(p)\pm\varepsilon$; the blue vertical strip is $p\pm\delta$. A $\delta$ works if the graph over the strip stays inside the band (drawn green). For $x^2$, shrink $\varepsilon$ and watch the largest working $\delta$ shrink with it, and notice it also shrinks as you move $p$ to the right. For the step at $p=0$ and for $\sin(1/x)$ at $p=0$, find an $\varepsilon$ for which no $\delta$ works. Then compare with $x\sin(1/x)$.

The widget's two discontinuities, written as sequence arguments:

  • Step function ($1$ for $x\ge0$, $0$ for $x\lt0$): $x_n=-\frac1n\to0$ but $f(x_n)=0\not\to1=f(0)$.
  • $\sin(1/x)$ with $f(0)=0$: $x_n=\frac1{\pi/2+2\pi n}\to0$ but $f(x_n)=1\not\to0$. No jump, just wild oscillation.
  • $x\sin(1/x)$ with $f(0)=0$ is continuous at 0: $|f(x)-0|\le|x|$, so $\delta=\varepsilon$ works.

Show that if $g:\R^n\to\R$ is continuous, then the sublevel set $S=\{\x: g(\x)\le c\}$ is closed. Deduce that $\mathcal F=\{\x\in\R^2 : x_1^2+x_2^2\le4,\ x_1+x_2\ge1\}$ is closed.

  1. Take any sequence $\x_k\in S$ with $\x_k\to\x$. We must show $\x\in S$.

    Sequential characterization of closed sets: limits of convergent sequences in $S$ must stay in $S$.

  2. Each $g(\x_k)\le c$. By continuity, $g(\x_k)\to g(\x)$.

    Sequential continuity.

  3. A limit of numbers that are all $\le c$ is $\le c$: if $g(\x)>c$, then with $\varepsilon=g(\x)-c$, eventually $g(\x_k)>g(\x)-\varepsilon=c$, a contradiction. So $g(\x)\le c$, i.e. $\x\in S$.

    This is where "$\le$" matters. With a strict "$\lt$" the limit could land exactly on $c$, outside the set, as $\frac1n\to0$ does for $\{x:x>0\}$.

  4. $\mathcal F=\{g_1\le4\}\cap\{g_2\le-1\}$ with $g_1=x_1^2+x_2^2$ and $g_2=-(x_1+x_2)$, both continuous. Each set is closed, and an intersection of closed sets is closed.

    This is the standard way to show a feasible set is closed: write every constraint as "continuous function $\le$ constant" (equalities are two such constraints).

The intermediate value theorem

Let $f$ be a continuous real function on $[a,b]$. If $f(a)\lt c\lt f(b)$, then there is a point $x\in(a,b)$ with $f(x)=c$. (Similarly if $f(a)>f(b)$.)

A continuous function can't get from one height to another without passing through every height in between. Its most common use: if $f(a)\lt0\lt f(b)$, then $f$ has a root in $(a,b)$. Halving the interval and keeping the half whose endpoints still have opposite signs is the bisection method, and the same "bracket then shrink" reasoning guarantees that the line searches of Part 6 can find an acceptable step. Both hypotheses matter: the step function above jumps from 0 to 1 and never takes the value $\frac12$.

Go deeper: why the IVT is true, and why its converse fails

Rudin proves 4.23 through connectedness. A set is connected if it can't be split into two nonempty pieces each of which stays away from the other's closure. The connected subsets of $\R$ are exactly the intervals ([Rudin] Thm. 2.47), and a continuous image of a connected set is connected ([Rudin] Thm. 4.22). So $f([a,b])$ is an interval containing $f(a)$ and $f(b)$, hence every $c$ between them.

The converse fails ([Rudin] Remark 4.24): a function can take every intermediate value on every interval and still be discontinuous. $\sin(1/x)$ with $f(0)=0$ is an example: on any interval around 0 it hits every value in $[-1,1]$.

For $a_n=\frac1{n^2}\to0$ and $\varepsilon=10^{-4}$, what is the smallest $N$ such that $|a_n|\lt\varepsilon$ for all $n\ge N$?

$\frac1{n^2}\lt10^{-4}\iff n^2>10^4$. Careful with the strict inequality at $n=100$.

$n^2>10^4\iff n>100$, so $N=101$. At $n=100$, $\frac1{n^2}=10^{-4}$ exactly, which is not strictly less than $\varepsilon$. Since $\frac1{n^2}$ decreases, all $n\ge101$ work.

For $f(x)=x^2$ at $p=1$ with $\varepsilon=0.21$, what is the largest $\delta$ such that $|x-1|\lt\delta$ implies $|x^2-1|\lt0.21$?

You need $0.79\lt x^2\lt1.21$, i.e. $\sqrt{0.79}\lt x\lt1.1$ for $x$ near 1. How far is each end from 1? The window $(1-\delta,1+\delta)$ must fit on both sides.

The safe $x$ near 1 form $(\sqrt{0.79},\,1.1)\approx(0.8888,\,1.1)$. The distances from 1 are $0.1$ (right) and $1-\sqrt{0.79}\approx0.1112$ (left). The symmetric window must fit both, so $\delta=\min(0.1,0.1112)=0.1$. The parabola is steeper on the right, which is the side that limits $\delta$.

$f(x)=x^3+x-1$ has $f(0)=-1\lt0\lt1=f(1)$, so by the IVT it has a root in $(0,1)$. Run three steps of bisection (each step: evaluate $f$ at the midpoint, keep the half whose endpoint values have opposite signs). What is the final interval?

$f(0.5)=-0.375$, so the root is in $[0.5,1]$. Next evaluate $f(0.75)$.

$f(0.5)=-0.375\lt0\Rightarrow[0.5,1]$. $f(0.75)=0.171875>0\Rightarrow[0.5,0.75]$. $f(0.625)\approx-0.1309\lt0\Rightarrow[0.625,0.75]$. (The root is about 0.6823.) Each step halves the interval, and the IVT guarantees a root stays inside.

Which of these functions on $\R$ is continuous at $x=0$?

For each one, try the sequences $x_n=\frac1{\pi/2+2\pi n}$ and $x_n=-\frac1n$. Which function is squeezed between $-|x|$ and $|x|$?

$|x\sin(1/x)|\le|x|$, so given $\varepsilon$, the choice $\delta=\varepsilon$ works: continuous. $\sin(1/x)$ along $x_n=\frac1{\pi/2+2\pi n}$ stays at 1, not 0. $1/x$ along $\frac1n$ blows up. The step along $-\frac1n$ stays at 0, not 1.

  • Simplify $|a_n-L|$ first, then solve "$\lt\varepsilon$" for $n$
  • Prove discontinuity with one bad sequence; prove continuity with $\varepsilon$–$\delta$ or the building-block rules
  • Show a feasible set is closed by writing it as $\{g_i\le c_i\}$ with continuous $g_i$
  • Check the IVT's sign change before bisecting
  • Choosing $N$ before $\varepsilon$ (the quantifier order matters)
  • Checking a few terms and declaring convergence
  • Assuming $\{g\lt c\}$ is closed
  • Using the IVT on a function with a jump
  1. $p_n\to p$: for every $\varepsilon>0$ there is $N$ with $d(p_n,p)\lt\varepsilon$ for all $n\ge N$ (Rudin 3.1).
  2. $f$ is continuous at $p$ when every $\varepsilon$ has a $\delta$ (Rudin 4.5), equivalently when $x_n\to p$ forces $f(x_n)\to f(p)$. Consequently $\{g\le c\}$ is closed for continuous $g$.
  3. IVT (Rudin 4.23): a continuous function on $[a,b]$ takes every value between $f(a)$ and $f(b)$. It underlies bisection and line-search bracketing.

In the definition of $p_n\to p$, the index $N$…

must be the same for every $\varepsilon$
That's the quantifiers swapped. For $\frac1n$, no single $N$ handles every $\varepsilon$.
may depend on $\varepsilon$
"For every $\varepsilon$ there is an $N$": $N$ is chosen after $\varepsilon$ is known.
must satisfy $d(p_N,p)=0$
Terms needn't ever equal the limit ($\frac1n\ne0$).

The best way to show a function is not continuous at $p$ is to…

check that it's not differentiable at $p$
$|x|$ isn't differentiable at 0 but is continuous there.
find one sequence $x_n\to p$ with $f(x_n)\not\to f(p)$
One bad sequence breaks sequential continuity, which is equivalent to $\varepsilon$–$\delta$ continuity.
show the graph has a corner
Corners are fine for continuity; it's jumps and wild oscillations that break it.

Which feasible set is guaranteed closed, given continuous $g$ and $h$?

$\{\x : g(\x)\lt0\}$
Strict inequalities give open sets: $\{x : x\lt0\}$ misses its limit point 0.
$\{\x : g(\x)\le0,\ h(\x)=0\}$
Each constraint is "continuous $\le$ constant" ($h=0$ is $h\le0$ and $-h\le0$), and intersections of closed sets are closed.
$\{\x : g(\x)\ne0\}$
That's the complement of the closed set $\{g=0\}$, so it's open, and typically not closed.

$f$ is continuous on $[0,2]$ with $f(0)=3$ and $f(2)=-1$. The IVT guarantees:

$f$ has a minimum at $x=2$
The IVT says nothing about where extremes are. (Weierstrass, next chapter, says a minimum exists somewhere.)
exactly one root in $(0,2)$
At least one; there may be several.
some $x\in(0,2)$ with $f(x)=1$
1 lies between $-1$ and $3$, so it's taken somewhere in between.

Compactness, the Weierstrass theorem, and coercivity: when a minimizer is guaranteed

A continuous function on a nonempty closed and bounded set in $\R^n$ always attains its minimum and its maximum; on an unbounded set, a function that grows to $+\infty$ in every direction (coercive) attains its minimum too.

This is the existence theorem behind every "let $\x^\star$ be the minimizer" in the course, and a classic exam question gives you $f$ and $\mathcal F$ and asks: does a minimizer exist, and why or why not?

Water poured into a bowl always finds a lowest point. Pour it onto a slope that falls forever (unbounded), or into a bowl with a pinhole at the bottom (not closed), and it runs away.

Compact sets

Chapter 1 showed that a closed and bounded set of numbers contains its inf and sup. The idea that makes this work in $\R^n$ is compactness. Its official definition uses open covers: a collection of open sets whose union contains $K$.

$K$ is compact if every open cover of $K$ has a finite subcover: whenever $K\subseteq\bigcup_\alpha G_\alpha$ with each $G_\alpha$ open, finitely many of the $G_\alpha$ already cover $K$.

A subset of $\R^n$ is compact if and only if it is closed and bounded.

Equivalently (Bolzano–Weierstrass): $K\subseteq\R^n$ is compact if and only if every sequence in $K$ has a subsequence converging to a point of $K$.

You'll almost never check the open-cover definition directly. In this course, "compact" means "closed and bounded" (the book's exam tip allows Heine–Borel as a black box). Examples: $[a,b]$, closed balls, closed boxes, the unit circle, a finite set, $\{\frac1n\}\cup\{0\}$. Not compact: $(0,1]$ (not closed), $[0,\infty)$ and $\Z$ (not bounded), $\{\frac1n\}$ (not closed).

If $f$ is a continuous mapping of a compact metric space $X$ into $\R^k$, then $f(X)$ is closed and bounded. In particular, $f$ is bounded on $X$.

Let $f$ be a continuous real function on a nonempty compact set $K$, with $M=\sup_{\x\in K}f(\x)$ and $m=\inf_{\x\in K}f(\x)$. Then there are points $\p,\mathbf q\in K$ with $f(\p)=M$ and $f(\mathbf q)=m$. That is, $f$ attains its maximum and its minimum on $K$.

How the pieces fit: by Theorem 4.15 the set of values $f(K)$ is closed and bounded; by Theorem 2.28 a closed, bounded set of reals contains its inf and sup. So the infimum value is some $f(\mathbf q)$: a minimizer. Memorize the hypotheses precisely: $f$ continuous, $K$ nonempty, closed and bounded. Drop any one and the conclusion can fail:

Hypothesis droppedCounterexampleWhat goes wrong
closed$f(x)=x$ on $(0,1]$$\inf=0$ at the missing endpoint: no minimum. ($f(x)=\frac1x$ on $(0,1]$ has no maximum.)
bounded$f(x)=x$ on $\R$; $f(x)=e^{-x}$ on $[0,\infty)$Unbounded below; or $\inf=0$ approached only as $x\to\infty$.
continuous$f(x)=x$ for $x>0$, $f(0)=1$, on $[0,1]$Values approach 0, but at 0 the function jumps up. (The book's example: $\frac1x$ with $f(0)=0$ on $[-1,1]$ has no max and no min.)
nonempty$K=\emptyset$Nothing to attain; by convention $\inf_\emptyset f=+\infty$.
Try it

Switch each hypothesis on or off. With all three on, Weierstrass guarantees a minimizer. Turn any off and you see a counterexample where the minimum is not attained: find the "escape route" each time (a missing endpoint, infinity, or a jump). Then switch to "lucky case": the same failed hypothesis, yet a minimizer exists anyway.

Careful: Weierstrass is a sufficient condition, not a necessary one. "$K$ isn't compact, so there's no minimizer" is a wrong argument: $x^2$ on $\R$ has a minimizer although $\R$ is unbounded. When a hypothesis fails, the theorem just stops giving a guarantee. To show a minimizer does not exist, compute the infimum and show no point attains it.

A classic exam mistake (the book's pitfall): "$f$ is continuous and $\mathcal F$ is closed, therefore a minimizer exists." Boundedness is not optional: $f(x)=x$ is continuous on the closed set $\R$ and has no minimizer. The book's three-step checklist:

  1. Is $f$ continuous on $\mathcal F$? (Nearly always, for the functions in this course.)
  2. Is $\mathcal F$ closed? Constraints of the form $g(\x)\le c$ or $h(\x)=c$ with continuous $g,h$ give closed sets; strict inequalities usually don't.
  3. Is $\mathcal F$ bounded? If not, is $f$ coercive? That's next.

Coercivity: Weierstrass on all of $\R^n$

Unconstrained problems live on $\R^n$, which is closed but not bounded, so Weierstrass doesn't apply directly. The fix: if $f$ grows without bound far away, the minimizer can't be far away, so we can restrict attention to a bounded region.

A function $f:\R^n\to\R$ is coercive if $f(\x)\to+\infty$ as $\norm\x\to\infty$: for every $M$ there is an $R$ such that $f(\x)>M$ whenever $\norm\x>R$.

The sublevel set at level $\alpha$ is $S_\alpha=\{\x: f(\x)\le\alpha\}$ (Part 0a). For coercive $f$, every sublevel set is bounded: taking $M=\alpha$, all of $S_\alpha$ lies in the ball of radius $R$.

If $f:\R^n\to\R$ is continuous and coercive, then $f$ has a global minimizer on $\R^n$. More generally, the same holds on any nonempty closed set $\mathcal F$ if $f(\x)\to+\infty$ as $\norm\x\to\infty$ within $\mathcal F$.

Prove the theorem. (The book says to be able to reproduce this chain verbatim: coercivity $\Rightarrow$ compact sublevel set $\Rightarrow$ Weierstrass $\Rightarrow$ minimizer.)

  1. Pick any point $\x_0$ and let $S=\{\x: f(\x)\le f(\x_0)\}$. Then $S$ is nonempty, because $\x_0\in S$.

    Choosing the level $f(\x_0)$ guarantees the set isn't empty, which Weierstrass needs.

  2. $S$ is closed: it's a sublevel set of a continuous function (Chapter 3).

    Limits of points with $f\le f(\x_0)$ also have $f\le f(\x_0)$.

  3. $S$ is bounded: by coercivity with $M=f(\x_0)$, there is $R$ with $f(\x)>f(\x_0)$ whenever $\norm\x>R$, so $S\subseteq\{\x:\norm\x\le R\}$.

    This is the only place coercivity is used: it fences the sublevel set into a ball.

  4. So $S$ is compact (Heine–Borel), and by Weierstrass $f$ attains its minimum on $S$ at some $\x^\star$.

    Now all three hypotheses hold, on $S$ instead of $\R^n$.

  5. $\x^\star$ is a global minimizer: for $\x\in S$, $f(\x^\star)\le f(\x)$ by construction; for $\x\notin S$, $f(\x)>f(\x_0)\ge f(\x^\star)$ because $\x_0\in S$.

    Points outside the sublevel set are worse than $\x_0$, which is no better than $\x^\star$. Nothing outside can win.

Try it

The green region is the sublevel set $\{f\le\alpha\}$; the dashed circle has radius $R$. The readout lists the minimum of $f$ on circles of growing radius: for a coercive function these minima grow without bound, so every sublevel set is fenced in. Raise $\alpha$ for the first two functions: the set grows but stays bounded. For $x^2-y^2$, $(x-y)^2$ and $e^x+y^2$, find the direction in which the sublevel set escapes.

(Tutorial 1, Problem 19.) Show that the non-convex function $h(x,y)=x^4-4xy+y^4$ has a global minimum, and find all points where it's attained.

  1. Let $r^2=x^2+y^2$. From $(x-y)^2\ge0$, $2xy\le x^2+y^2$, so $-4xy\ge-2r^2$.

    Bound the troublesome cross term by something depending only on the distance $r$ from the origin.

  2. From $(x^2-y^2)^2\ge0$, $2x^2y^2\le x^4+y^4$, so $r^4=x^4+2x^2y^2+y^4\le2(x^4+y^4)$, i.e. $x^4+y^4\ge\frac{r^4}2$.

    Same trick one degree up: the quartic part is at least half of $r^4$.

  3. So $h\ge\frac{r^4}2-2r^2\to+\infty$ as $r\to\infty$: $h$ is coercive. It's a polynomial, hence continuous, so a global minimizer exists.

    The theorem just proved. Note that $h$ isn't convex, so existence here comes from coercivity alone.

  4. A global minimizer of a smooth function on $\R^2$ is a stationary point (Part 3): $\grad h=(4x^3-4y,\ 4y^3-4x)=\0$, so $y=x^3$ and $x=y^3=x^9$. Then $x(x^8-1)=0$, giving $(0,0),(1,1),(-1,-1)$.

    Existence first, then search: since a minimizer exists, it must be one of these candidates.

  5. $h(0,0)=0$ and $h(\pm1,\pm1)=1-4+1=-2$. The global minimum is $-2$, attained at $(1,1)$ and $(-1,-1)$. (The origin is a saddle point.)

    Compare values. Without step 3, the smallest stationary value need not be the global minimum: the infimum could be $-\infty$ or unattained.

The running example, and where this goes next

For a quadratic $f(\x)=\frac12\x^\top A\x-\b^\top\x$ with $A\succ0$ and smallest eigenvalue $m>0$ (Part 0b), $\x^\top A\x\ge m\norm\x^2$ and, by Cauchy–Schwarz, $\b^\top\x\le\norm\b\norm\x$. So $$f(\x)\ \ge\ \tfrac m2\norm\x^2-\norm\b\,\norm\x\ \to\ +\infty:$$ positive definite quadratics are coercive, and always have a minimizer. Our fit-the-line loss $L(w,c)=\frac12\z^\top A\z-\b^\top\z+38$ has $m\approx0.721$, so it's coercive, and its minimizer $(1.5,\frac13)$ was guaranteed to exist before we computed it.

The same argument with the strong-convexity inequality in place of $\frac m2\norm\x^2$ is Tutorial 1, Problem 17(b): a strongly convex function on $\R^n$ has exactly one global minimizer (existence by coercivity and Weierstrass; uniqueness by strict convexity). Part 4 does it in full. Coercivity can also come from constraints: in the diet problem of Part 1, on the feasible set $\x\ge\0$, the cost satisfies $0.30x_1+0.25x_2\ge0.25\norm\x_1\to\infty$, so a cheapest diet exists.

Go deeper: proofs of Theorems 4.14–4.16 and Heine–Borel

4.14 (compact image). Let $\{V_\alpha\}$ be an open cover of $f(X)$. Continuity makes each preimage $f^{-1}(V_\alpha)$ open ([Rudin] Thm. 4.8), and they cover $X$. Compactness gives finitely many covering $X$; their $V_\alpha$ cover $f(X)$. So $f(X)$ is compact, hence closed and bounded (4.15, via [Rudin] Thm. 2.41).

4.16 (Weierstrass). $f(X)\subseteq\R$ is closed and bounded, so by 2.28 its sup $M$ and inf $m$ belong to it: $M=f(\p)$ and $m=f(\mathbf q)$ for some $\p,\mathbf q$.

A sequence proof. Take $\x_k\in K$ with $f(\x_k)\to m$ (possible by the definition of inf; if $m=-\infty$, with $f(\x_k)\to-\infty$). By Bolzano–Weierstrass a subsequence converges to some $\mathbf q\in K$ (closed). By continuity $f(\mathbf q)=\lim f(\x_{k_j})=m$, which also shows $m$ is finite.

Heine–Borel (sketch). Compact $\Rightarrow$ bounded: the balls $N_k(\0)$, $k=1,2,\dots$ cover $K$, and finitely many suffice. Compact $\Rightarrow$ closed: a point outside $K$ can be separated from $K$ by finitely many balls. Closed and bounded $\Rightarrow$ compact: if a box had an open cover with no finite subcover, bisect it repeatedly, always keeping a piece with no finite subcover. The nested pieces shrink to a point (completeness again), which lies in one open set of the cover, and that set swallows a small enough piece: contradiction.

Does $f(x,y)=\dfrac1{1+x^2+y^2}$ attain a minimum on $\R^2$?

What happens to $f$ far from the origin? Can $f$ ever equal that value? (Does it attain a maximum?)

$f>0$ everywhere and $f\to0$ as $x^2+y^2\to\infty$, so $\inf f=0$, but no point gives 0: no minimizer. Weierstrass doesn't apply because $\R^2$ isn't bounded, and $f$ is the opposite of coercive. The maximum is attained: $f\le1=f(0,0)$.

Show that $g(x,y)=x^4+y^4-32xy$ is coercive and find its global minimum value.

As in Problem 19: $-32xy\ge-16r^2$ and $x^4+y^4\ge\frac{r^4}2$. For the stationary points, $\grad g=\0$ gives $y=\frac{x^3}8$ and $x=\frac{y^3}8$.

$g\ge\frac{r^4}2-16r^2\to\infty$, so $g$ is coercive and continuous: a global minimizer exists and is a stationary point. $4x^3=32y$ and $4y^3=32x$ give $y=x^3/8$, $x=y^3/8=x^9/8^4$, so $x=0$ or $x^8=4096$, i.e. $x=\pm2\sqrt2$ (since $(2\sqrt2)^8=2^{12}$). Then $y=\frac{(2\sqrt2)^3}{8}=2\sqrt2$ with the same sign. $g(\pm2\sqrt2,\pm2\sqrt2)=64+64-32\cdot8=-128$, versus $g(0,0)=0$. The global minimum is $-128$.

Which function is coercive on $\R^2$?

For each, look for a direction (a line or ray to infinity) along which the function stays bounded or goes to $-\infty$.

$x^4+y^2-3y=x^4+(y-\frac32)^2-\frac94\to\infty$ in every direction: coercive. The others fail: $(x+y)^2=0$ along $y=-x$; $x^2+y^2-10xy=-8x^2$ along $y=x$; $e^x+e^y\to0$ as $x,y\to-\infty$.

For the fit-the-line quadratic $f(\z)=\frac12\z^\top A\z-\b^\top\z$ with $A=\begin{pmatrix}28&12\\12&6\end{pmatrix}$ and $\b=(46,20)$, use $f(\z)\ge\frac m2\norm\z^2-\norm\b\norm\z$ to find the radius $R$ such that $f(\z)>0=f(\0)$ whenever $\norm\z>R$. (So the minimizer lies in the ball of radius $R$.) Give $R$ to 3 significant figures.

$\frac m2R^2-\norm\b R>0\iff R>\frac{2\norm\b}m$. The eigenvalues of $A$ are $17\pm\sqrt{265}$; $\norm\b=\sqrt{46^2+20^2}$.

$m=17-\sqrt{265}=\frac{34-\sqrt{1060}}2\approx0.7212$ and $\norm\b=\sqrt{2516}\approx50.16$. So $R=\frac{2\norm\b}m\approx139$ (more precisely 139.10). The bound is crude (the minimizer $(1.5,\frac13)$ has norm about 1.54) because it uses only the flattest direction, but crude bounds are all existence needs.

  • Run the checklist in order: continuous? closed? bounded (or coercive)?
  • For unconstrained problems, prove coercivity with a lower bound $f\ge\phi(\norm\x)$ where $\phi(r)\to\infty$
  • Establish existence first, then find the minimizer among the stationary points
  • To show no minimizer exists, compute $\inf f$ and show it isn't attained
  • "Closed and continuous, so a minimizer exists" (forgetting boundedness)
  • "Not compact, so no minimizer" (Weierstrass is only sufficient)
  • Calling $f$ coercive because it's bounded below, as with $e^x+y^2$
  • Taking the smallest stationary value as the global minimum without proving a minimizer exists
  1. Heine–Borel: in $\R^n$, compact means closed and bounded.
  2. Weierstrass (Rudin 4.16): $f$ continuous on a nonempty compact $K$ attains its max and min. Drop closed, bounded or continuous and it can fail; the conditions are sufficient, not necessary.
  3. Coercive + continuous on $\R^n$ (or on a closed $\mathcal F$) gives a global minimizer, via a compact sublevel set. Positive definite quadratics and Tutorial 1's $x^4-4xy+y^4$ are coercive.

Which set is compact in $\R^2$?

$\{(x,y): x^2+y^2\lt1\}$
Bounded but not closed: the circle is missing.
$\{(x,y): y\ge x^2\}$
Closed but not bounded.
$\{(x,y): x^2+y^2=1\}$
The unit circle is closed ($g=x^2+y^2$ is continuous and it's $\{g=1\}$) and bounded. Heine–Borel says compact.
$\{(\frac1n,0): n\in\N\}$
Bounded, but its limit point $(0,0)$ is missing.

A student argues: "$\min_{x\in\R}x^2$ has no solution because $\R$ is not compact." The flaw is:

$\R$ is compact
$\R$ is closed but unbounded, so not compact.
Weierstrass gives a sufficient condition, not a necessary one
When compactness fails the theorem is silent. Here $x^\star=0$ exists, and coercivity of $x^2$ proves it.
$x^2$ is not continuous
Polynomials are continuous.

In the proof that coercive continuous $f$ has a minimizer, why use the sublevel set at level $f(\x_0)$ rather than any level?

Because it's open
Sublevel sets $\{f\le\alpha\}$ of continuous functions are closed, not open.
Because it's unbounded
Coercivity makes it bounded; that's the point.
It contains $\x_0$, so it's nonempty, and everything outside it is worse than $\x_0$
Nonempty is needed for Weierstrass, and "worse than $\x_0$ outside" turns the minimizer on $S$ into a global one.

Which pair of hypotheses already guarantees that $\min_{\x\in\mathcal F}f(\x)$ has a solution, for $f$ continuous on $\R^n$?

$\mathcal F$ closed and $f$ bounded below
$e^{x}$ on $\R$: closed set, bounded below, no minimizer.
$\mathcal F$ nonempty and closed, and $f$ coercive
The sublevel set within $\mathcal F$ is closed and bounded, hence compact; Weierstrass does the rest.
$\mathcal F$ bounded and nonempty
$(0,1]$ with $f(x)=x$: bounded, but not closed.

$f(\x)=\frac12\x^\top A\x-\b^\top\x$ with $A$ symmetric and $\lambda_{\min}(A)>0$. Then:

a minimizer may fail to exist if $\b$ is large
Large $\b$ only moves the minimizer. The quadratic term eventually dominates the linear one.
$f$ is coercive only if $\b=\0$
$f\ge\frac m2\norm\x^2-\norm\b\norm\x$ grows to infinity for any $\b$.
$f$ is coercive, so a global minimizer exists
Positive definite quadratics are coercive; Part 1 noted the minimizer is $A^{-1}\b$.

$E=\{\frac{(-1)^n}n : n\in\N\}$. Which is correct?

$\sup E=0$, not attained
0 is the limit of the terms. Is it an upper bound? What is the term at $n=2$?
$\sup E=\frac12$ and $\inf E=-1$, both attained
The largest term is $\frac12$ ($n=2$) and the smallest is $-1$ ($n=1$).
$\sup E=1$, $\inf E=-1$
No term equals or approaches 1. List the first few terms.

Gradient descent on $f(x)=e^{-x}$ over $[0,\infty)$ produces iterates that run off to $+\infty$. The best explanation is:

the step size is too large
Any step size has the same problem: there's nothing to converge to.
$\inf f=0$ is not attained: there is no minimizer, because $[0,\infty)$ is unbounded and $f$ isn't coercive
The infimum escapes to infinity, so the iterates chase it there.
$f$ is not continuous
$e^{-x}$ is continuous.

The punctured disc $\{\x\in\R^2: 0\lt\norm\x\lt1\}$ is:

closed, since it excludes the centre
The centre is a limit point that's excluded: that makes it not closed.
open, not closed; its closure is the closed unit disc
Every point has a ball inside (avoiding the centre and the rim); the centre and rim are missing limit points.
neither open nor closed
Every point of the set is interior: shrink the ball below both the distance to the centre and the distance to the rim.

To prove $\{\x\in\R^n : \norm{\x}\le5,\ x_1\ge0\}$ is closed, the cleanest argument is:

it's bounded, so it's closed
Bounded and closed are independent: $(0,1)$ is bounded and not closed.
it's an intersection of sublevel sets $\{\norm\x\le5\}$ and $\{-x_1\le0\}$ of continuous functions
Each is closed (limits of sequences stay inside), and intersections of closed sets are closed.
its complement is bounded
Its complement is unbounded. The correct link is: closed exactly when the complement is open.

$h(x,y)=x^4-4xy+y^4$ is not convex. Why does it still have a global minimizer?

Because $\grad h=\0$ has solutions
Stationary points exist for $x^2-y^2$ too, which has no minimizer.
Because $h\ge\frac{r^4}2-2r^2$ makes it coercive, and it's continuous
Coercivity gives a compact sublevel set; Weierstrass gives the minimizer, which turns out to be at $\pm(1,1)$ with value $-2$.
Because its Hessian is positive definite everywhere
At the origin the Hessian has eigenvalues $\pm4$: indefinite.

For the fit-the-line loss, which fact about $A=\begin{pmatrix}28&12\\12&6\end{pmatrix}$ makes $L$ coercive?

Its entries are positive
Positive entries don't imply positive definiteness, and that's not the right property anyway.
Its determinant is 24
A positive determinant alone allows both eigenvalues negative. You need both positive.
Its smallest eigenvalue $m\approx0.721$ is positive, so $\z^\top A\z\ge m\norm\z^2$
Then $L\ge\frac m2\norm\z^2-\norm\b\norm\z+38\to\infty$.

Lectures 4 and 5 (18 and 20 Aug): how to recognise a local minimum from derivatives alone. The first-order condition $\grad f(\x^\star)=\0$ narrows the search to critical points; the Hessian's eigenvalues then sort them into minima, maxima and saddles, except in the degenerate cases where the Hessian says nothing. You'll prove all three conditions (FONC, SONC, SOSC) from Taylor's theorem, which is exactly what the exam asks.

You need: Part 0b (gradient and directional derivative, eigenvalues and definiteness, the Hessian and the three forms of Taylor's theorem) and Part 1 (local vs global minimizers).

Critical points and the first-order necessary condition

At a local minimum of a smooth function, the slope in every direction is zero: $\grad f(\x^\star)=\0$.

This turns "search all of $\R^n$" into "solve $n$ equations", and its proof (a one-variable argument along a line) is a standard exam question.

Standing at the bottom of a bowl in the dark: if the ground under your feet tilted down in any direction, you could step that way and go lower, so you weren't at the bottom.

The setting

From now on the problem is unconstrained: $$\min_{\x\in\R^n} f(\x).$$ More generally, $f$ may be defined only on an open set $U\subseteq\R^n$: a set where every point has a small ball around it that stays inside $U$. Openness is what lets you step a little in every direction from any point, and every proof in this part uses that.

How smooth is smooth enough?

The theorems carry hypotheses about how many derivatives $f$ has. Recall from Part 0b:

  • $f\in C^1$ ("continuously differentiable") if all first partial derivatives exist and are continuous. Then $\grad f$ exists and the directional derivative is $D_{\d}f(\x)=\grad f(\x)^\top\d$.
  • $f\in C^2$ if all second partial derivatives exist and are continuous. Then the Hessian $\hess f(\x)$ exists, is symmetric (Schwarz's theorem), and so has real eigenvalues and perpendicular eigenvectors (spectral theorem).

The hypotheses are not decoration. $f(x)=|x|$ is not even differentiable at 0. $f(x)=x|x|$ has $f'(x)=2|x|$, which is continuous, so $f\in C^1$; but $f'$ has a kink at 0, so $f''(0)$ doesn't exist and $f\notin C^2$. The first-order theory below needs $C^1$; the second-order theory (Chapter 3.2) needs $C^2$, both to have a Hessian at all and to know it's symmetric.

A point $\x^\star$ with $\grad f(\x^\star)=\0$ is a critical point (also called a stationary point) of $f$.

Careful: Tutorial 1 (Def. 4) and the book's chapter for Lecture 4 also count points where $f$ is not differentiable as critical points (useful for functions like $|x|$). Both then immediately restrict to smooth $f$, as the lectures do. In this course, "critical point" means $\grad f(\x^\star)=\0$.

Let $f\in C^1$ on an open set $U\subseteq\R^n$. If $\x^\star\in U$ is a local minimizer (or local maximizer) of $f$, then $$\grad f(\x^\star)=\0.$$

The idea of the proof: look at $f$ only along a straight line through $\x^\star$. That's a function of one variable, it has a local minimum at the centre, so its derivative there is zero. Since that's true for every line, the gradient must be zero. Learn the steps below; you should be able to write them out from memory.

Prove the FONC.

  1. Let $\x^\star$ be a local minimizer: $f(\x^\star)\le f(\x)$ for all $\x$ with $\norm{\x-\x^\star}\lt\varepsilon$. Fix any direction $\d\ne\0$ and define the one-variable function $g(t)=f(\x^\star+t\d)$.

    Since $U$ is open, $\x^\star+t\d\in U$ for all small $|t|$, so $g$ is defined near $t=0$. We have turned an $n$-variable problem into a 1-variable one.

  2. For $|t|\lt\varepsilon/\norm{\d}$, the point $\x^\star+t\d$ lies in the $\varepsilon$-ball, so $g(t)\ge g(0)$: $t=0$ is a local minimizer of $g$.

    The distance from $\x^\star$ is $|t|\,\norm{\d}\lt\varepsilon$.

  3. For small $t\gt0$: $\dfrac{g(t)-g(0)}{t}\ge0$ (numerator $\ge0$, denominator $\gt0$). Letting $t\to0^+$ gives $g'(0)\ge0$.

    A limit of nonnegative numbers is nonnegative.

  4. For small $t\lt0$: $\dfrac{g(t)-g(0)}{t}\le0$ (numerator $\ge0$, denominator $\lt0$). Letting $t\to0^-$ gives $g'(0)\le0$. Hence $g'(0)=0$.

    $g$ is differentiable (chain rule, since $f\in C^1$), so both one-sided limits equal the same number $g'(0)$. It is both $\ge0$ and $\le0$.

  5. By the chain rule, $g'(0)=\grad f(\x^\star)^\top\d$. So $\grad f(\x^\star)^\top\d=0$ for every $\d$.

    This is the directional-derivative formula from Part 0b.

  6. Choose $\d=\grad f(\x^\star)$ (if it were nonzero): then $\norm{\grad f(\x^\star)}^2=0$, so $\grad f(\x^\star)=\0$. For a local maximizer, apply the same argument to $-f$. $\blacksquare$

    A vector whose dot product with every vector is zero (in particular with itself) must be the zero vector.

Go deeper: the Taylor proof, and Luenberger's version with constraints

Via Taylor. The asymptotic form gives $f(\x^\star+t\d)-f(\x^\star)=t\,\grad f(\x^\star)^\top\d+o(t)$. If $\grad f(\x^\star)^\top\d\ne0$ for some $\d$, the linear term dominates for small $t$ and changes sign as $t$ crosses 0, so $f$ is lower than $f(\x^\star)$ on one side: contradiction. Either proof is accepted on the exam.

[LD] §7.1 (p.184–185). Luenberger proves a more general statement. Call $\d$ a feasible direction at $\x^\star$ in a set $\Omega$ if $\x^\star+t\d\in\Omega$ for all small $t\gt0$. If $\x^\star$ is a local minimizer over $\Omega$, then $\grad f(\x^\star)^\top\d\ge0$ for every feasible direction. Only step 3 above survives (we can only move forward, $t\gt0$), so we get an inequality instead of an equality. At an interior point both $\d$ and $-\d$ are feasible, which gives back $\grad f(\x^\star)=\0$. At a boundary point, only the inequality holds: that's why the diet LP of Part 1 had its minimum at a corner with $\grad f\ne\0$.

Careful: FONC is necessary, not sufficient. $f(x)=x^3$ has $f'(0)=0$, yet 0 is not a minimum or maximum (it's an inflection point). $f(x)=x^4-2x^2$ has $f'(0)=0$ at a local maximum; Part 1 showed gradient descent getting stuck there. And in two or more dimensions there are saddle points such as the origin for $x^2-y^2$. FONC only produces candidates.
Try it

Drag the black point and turn the direction $\d$. The small plot is the one-variable slice $g(t)=f(\x+t\d)$ used in the proof, with its tangent at $t=0$ (slope $g'(0)=\grad f^\top\d$). Wherever $\grad f\ne\0$, press "Point d along −∇f" and see the slice go downhill. Then snap to a critical point and turn $\d$ through all angles: every slice is flat at $t=0$. On the saddle, notice that some flat slices still go down on both sides.

Find all critical points of $f(x,y)=x^3-3x+y^3-3y^2$ (Tutorial 1, Problem 4).

  1. $\grad f(x,y)=\begin{pmatrix}3x^2-3\\3y^2-6y\end{pmatrix}$.

    Differentiate in $x$ holding $y$ fixed, then in $y$ holding $x$ fixed.

  2. $3x^2-3=0\Rightarrow x=\pm1$. $\ 3y^2-6y=3y(y-2)=0\Rightarrow y\in\{0,2\}$.

    Here the two equations are decoupled: each involves only one variable, so solve them separately.

  3. Critical points: $(1,0)$, $(1,2)$, $(-1,0)$, $(-1,2)$, with values $f=-2,\ -6,\ 2,\ -2$.

    Every combination of an $x$-root with a $y$-root. FONC alone can't say which, if any, is a minimum; Chapter 3.2 finishes the job.

Running example: fit the line

For the loss $L(\z)=\tfrac12\z^\top A\z-\b^\top\z+38$ with $A=\begin{pmatrix}28&12\\12&6\end{pmatrix}$, $\b=(46,20)^\top$, FONC says $\grad L=A\z-\b=\0$, i.e. the linear system $28w+12c=46$, $12w+6c=20$. Subtracting twice the second equation from the first gives $4w=6$, so $w=1.5$ and $c=(20-18)/6=1/3$. These are the normal equations of least squares: fitting a model by "setting the derivative to zero" is FONC. (That this critical point is the global minimizer needs more: Chapter 3.2 and Part 4.)

([LD] §7.1, Example 1.) Find the critical point of $f(x_1,x_2)=x_1^2-x_1x_2+x_2^2-3x_2$.

Set both partial derivatives to zero: $2x_1-x_2=0$ and $-x_1+2x_2-3=0$.

From the first equation $x_2=2x_1$. Substituting: $-x_1+4x_1=3$, so $x_1=1$, $x_2=2$. (It is in fact the global minimizer: the Hessian $\begin{pmatrix}2&-1\\-1&2\end{pmatrix}$ has eigenvalues 1 and 3.)

How many critical points does $f(x,y)=x^4+y^4-4xy$ have?

$\grad f=(4x^3-4y,\ 4y^3-4x)$. The equations are coupled: the first gives $y=x^3$; substitute into the second.

$y=x^3$ and $x=y^3$ give $x=x^9$, i.e. $x(x^8-1)=0$, so $x\in\{0,1,-1\}$ (real roots only), with $y=x^3$. Three critical points: $(0,0)$, $(1,1)$, $(-1,-1)$.

Classify $f(x)=x|x|$ on $\R$.

Write $f(x)=x^2$ for $x\ge0$ and $-x^2$ for $x\lt0$. Compute $f'$ on each side and at 0. Is $f'$ continuous? Is $f'$ differentiable at 0?

$f'(x)=2x$ for $x\gt0$, $-2x$ for $x\lt0$, and $f'(0)=\lim_{h\to0}h|h|/h=0$. So $f'(x)=2|x|$: continuous, hence $f\in C^1$. But $2|x|$ has a kink at 0, so $f''(0)$ does not exist: $f\notin C^2$.

At a point $\x$, $\grad f(\x)=(3,4)^\top$. As in the FONC proof, let $g(t)=f(\x+t\d)$ with $\d=-\grad f(\x)$. What is $g'(0)$? (This is the number that shows $\x$ cannot be a local minimizer.)

$g'(0)=\grad f(\x)^\top\d$.

$g'(0)=(3,4)\cdot(-3,-4)=-25=-\norm{\grad f}^2\lt0$. So $g(t)\lt g(0)$ for small $t\gt0$: moving along $-\grad f$ decreases $f$, and $\x$ is not a local minimizer.

  • Reduce to one variable with $g(t)=f(\x^\star+t\d)$ and use one-sided limits
  • State the hypotheses: $f\in C^1$, $\x^\star$ an interior point of an open set
  • Treat solutions of $\grad f=\0$ as candidates only
  • When the equations are coupled, substitute one into the other and check every case
  • Concluding "minimum" from $\grad f=\0$ alone ($x^3$, saddles)
  • Applying $\grad f=\0$ at a boundary point of a constraint set
  • Forgetting to divide by a negative $t$ flips the inequality in step 4
  • Dropping solutions when factoring, e.g. losing $x=0$ in $x(x^8-1)=0$
  1. FONC: if $f\in C^1$ and $\x^\star$ is an interior local minimizer or maximizer, then $\grad f(\x^\star)=\0$.
  2. Proof: the slice $g(t)=f(\x^\star+t\d)$ has a local min at 0, so $g'(0)=\grad f(\x^\star)^\top\d=0$ for every $\d$; take $\d=\grad f(\x^\star)$.
  3. Critical points are only candidates: minima, maxima, inflections and saddles all have $\grad f=\0$.

Which statement is true for $f\in C^1(\R^n)$?

If $\grad f(\x^\star)=\0$ then $\x^\star$ is a local minimizer or maximizer
Think of $x^3$ at 0, or $x^2-y^2$ at the origin.
If $\x^\star$ is a local maximizer then $\grad f(\x^\star)=\0$
FONC covers maxima too: apply the minimum case to $-f$.
If $\grad f(\x^\star)\ne\0$, $\x^\star$ may still be a local minimizer of $f$ on $\R^n$
On all of $\R^n$ (an open set) a local minimizer must have zero gradient. Only on a constrained set can a minimizer have $\grad f\ne\0$.

In the FONC proof, where is it used that $U$ is open?

To make $g$ differentiable
Differentiability of $g$ comes from $f\in C^1$ and the chain rule.
To choose $\d=\grad f(\x^\star)$
That choice is always allowed; it needs no property of $U$.
To make $g(t)=f(\x^\star+t\d)$ defined for small negative and positive $t$
Both one-sided limits are needed. At a boundary point you may only be able to move one way, and then you get just $g'(0)\ge0$.

In step 4 of the proof ($t\lt0$), why does the difference quotient satisfy $\frac{g(t)-g(0)}{t}\le0$?

Because $g$ is decreasing to the left of 0
We don't know $g$ is monotone. Look at the signs of numerator and denominator separately.
The numerator is $\ge0$ (local min) and the denominator is negative
A nonnegative number divided by a negative one is $\le0$.
Because $g'(0)=0$
That's the conclusion we're proving; using it here would be circular.

$f(x,y)=x^2+y^2$ restricted to the closed half-plane $x\ge1$ has its minimizer at $(1,0)$. What is $\grad f(1,0)$?

$(0,0)$, by FONC
FONC needs an interior point. Is $(1,0)$ in the interior of $\{x\ge1\}$?
$(2,0)$: nonzero, because the minimizer is on the boundary
Every feasible direction $\d$ has $d_1\ge0$, and indeed $\grad f^\top\d=2d_1\ge0$: Luenberger's one-sided condition.
Undefined, because the set is not open
$f$ is a polynomial; its gradient exists everywhere.

Second-order conditions: telling minima, maxima and saddles apart

At a critical point the gradient is zero, so the Hessian takes over: all eigenvalues positive means a strict local minimum, all negative a strict local maximum, mixed signs a saddle.

"Find and classify all critical points" is the most common exam question on this material, and the proofs of SONC and SOSC are examinable.

A ball resting on a surface: in a bowl it rolls back after a nudge (minimum); on a hilltop it rolls off any way (maximum); on a horse's saddle it rolls off sideways but back along the horse's spine (saddle).

What Taylor says at a critical point

At a critical point, $\grad f(\x^\star)=\0$, so the asymptotic Taylor formula (Part 0b) loses its linear term: $$f(\x^\star+\p)-f(\x^\star)=\tfrac12\,\p^\top\hess f(\x^\star)\,\p+o(\norm{\p}^2).$$ For small steps, the quadratic form decides. Write $H=\hess f(\x^\star)$, with eigenvalues $\lambda_i$ and orthonormal eigenvectors $\u_i$, and expand $\p=\sum_ic_i\u_i$. Then, as in Part 0b, $$\tfrac12\p^\top H\p=\tfrac12\left(\lambda_1c_1^2+\dots+\lambda_nc_n^2\right).$$ If every $\lambda_i\gt0$, every small step goes up: a bowl. If every $\lambda_i\lt0$, every step goes down: a cap. If some are positive and some negative, you go up along some eigenvectors and down along others: a saddle. If some $\lambda_i=0$, the quadratic part is flat in that direction, and the $o(\norm{\p}^2)$ terms we threw away decide.

$f_1(x)=x^4$, $f_2(x)=-x^4$ and $f_3(x)=x^3$ all have $f'(0)=0$ and $f''(0)=0$. At $x=0$…

all three have a minimum, since $f''(0)\ge0$
$f''(0)\ge0$ is necessary for a minimum, not sufficient. Look at $-x^4$.
the second-derivative test tells them apart
It sees the same numbers, $f'(0)=f''(0)=0$, for all three.
$f_1$ has a minimum, $f_2$ a maximum, $f_3$ neither, and second derivatives can't tell which
When the second derivative vanishes, higher-order terms decide. This is the "degenerate" case of Chapter 3.3.

Let $f\in C^2$ on an open set $U$. If $\x^\star\in U$ is a local minimizer, then $\grad f(\x^\star)=\0$ and $\hess f(\x^\star)\succeq0$ (positive semidefinite).

Prove SONC. (Tool: Taylor with Lagrange remainder, the standard tool for necessary conditions.)

  1. $\grad f(\x^\star)=\0$ is FONC. Now fix any $\d\in\R^n$. For small $t\gt0$, Taylor's Lagrange form gives some $\theta_t\in(0,1)$ with $$f(\x^\star+t\d)=f(\x^\star)+t\,\grad f(\x^\star)^\top\d+\tfrac{t^2}{2}\,\d^\top\hess f(\x^\star+\theta_tt\d)\,\d.$$

    Lagrange form is exact, at the price of evaluating the Hessian at an unknown point between $\x^\star$ and $\x^\star+t\d$.

  2. The gradient term is zero, so $f(\x^\star+t\d)-f(\x^\star)=\tfrac{t^2}{2}\,\d^\top\hess f(\x^\star+\theta_tt\d)\,\d$.

    Use FONC.

  3. Local minimality makes the left side $\ge0$ for small $t$. Divide by $t^2/2\gt0$: $\ \d^\top\hess f(\x^\star+\theta_tt\d)\,\d\ge0$.

    Dividing by a positive number keeps the inequality.

  4. Let $t\to0^+$. Since $0\lt\theta_t\lt1$, the point $\x^\star+\theta_tt\d\to\x^\star$, and $\hess f$ is continuous ($f\in C^2$), so $\d^\top\hess f(\x^\star)\,\d\ge0$.

    This is where $C^2$ (continuity of the second derivatives) is used. A limit of nonnegative numbers is nonnegative.

  5. $\d$ was arbitrary, so $\hess f(\x^\star)\succeq0$. $\blacksquare$

    That is the definition of positive semidefinite.

Let $f\in C^2$ on an open set $U$ and $\x^\star\in U$. If $\grad f(\x^\star)=\0$ and $\hess f(\x^\star)\succ0$ (positive definite), then $\x^\star$ is a strict local minimizer. If instead $\hess f(\x^\star)\prec0$, it is a strict local maximizer.

Prove SOSC. (Tool: the asymptotic Taylor form, the standard tool for sufficient conditions, plus the Rayleigh bound.)

  1. By the asymptotic form and $\grad f(\x^\star)=\0$: $$f(\x^\star+\p)-f(\x^\star)=\tfrac12\p^\top\hess f(\x^\star)\p+o(\norm{\p}^2).$$

    Here the Hessian is evaluated at $\x^\star$ itself, which is what lets us use its eigenvalues.

  2. Let $\lambda_{\min}\gt0$ be the smallest eigenvalue of $\hess f(\x^\star)$. By the Rayleigh bound, $\p^\top\hess f(\x^\star)\p\ge\lambda_{\min}\norm{\p}^2$ for all $\p$.

    Positive definite means every eigenvalue is positive, so $\lambda_{\min}\gt0$.

  3. Hence, for $\p\ne\0$, $$f(\x^\star+\p)-f(\x^\star)\ge\norm{\p}^2\left(\tfrac12\lambda_{\min}+\frac{o(\norm{\p}^2)}{\norm{\p}^2}\right).$$

    Factor out $\norm{\p}^2$ to compare the two terms.

  4. By definition of little-o, $o(\norm{\p}^2)/\norm{\p}^2\to0$, so there's an $\varepsilon\gt0$ with $\big|o(\norm{\p}^2)/\norm{\p}^2\big|\le\tfrac14\lambda_{\min}$ whenever $0\lt\norm{\p}\lt\varepsilon$. Then the bracket is $\ge\tfrac14\lambda_{\min}\gt0$.

    "Small enough steps" made precise: the error eventually drops below any fixed fraction of the curvature term.

  5. So $f(\x^\star+\p)\gt f(\x^\star)$ for all $0\lt\norm{\p}\lt\varepsilon$: a strict local minimizer. For $\hess f\prec0$, apply this to $-f$. $\blacksquare$

    Strict: every other point in the ball is strictly higher.

Careful: the gap between SONC and SOSC. SONC says minimum ⇒ $H\succeq0$; SOSC says $H\succ0$ ⇒ strict minimum. Neither converse holds. $x^4$ has a strict minimum at 0 with $f''(0)=0$, so $H\succ0$ is not necessary. $x^3$ satisfies $f'(0)=0$, $f''(0)\ge0$ and has no minimum, so $H\succeq0$ is not sufficient. When $H$ is semidefinite and singular, the second-order test is silent.

A critical point $\x^\star$ is a saddle point if every neighbourhood of $\x^\star$ contains points $\y,\z$ with $f(\y)\lt f(\x^\star)\lt f(\z)$.

Let $f\in C^2$, $\grad f(\x^\star)=\0$, and let $\lambda_1\le\dots\le\lambda_n$ be the eigenvalues of $\hess f(\x^\star)$.

  1. All $\lambda_i\gt0$ ($H\succ0$): strict local minimum.
  2. All $\lambda_i\lt0$ ($H\prec0$): strict local maximum.
  3. Some $\lambda_i\gt0$ and some $\lambda_j\lt0$ ($H$ indefinite): saddle point.
  4. Some $\lambda_i=0$ and no eigenvalues of opposite signs ($H$ semidefinite and singular): inconclusive; higher-order terms decide.

Cases 1 and 2 are SOSC. Proof of case 3. Take unit eigenvectors $\u_+$ with $\lambda_+\gt0$ and $\u_-$ with $\lambda_-\lt0$. With $\p=\alpha\u_+$, Taylor gives $f(\x^\star+\alpha\u_+)-f(\x^\star)=\tfrac12\alpha^2\lambda_++o(\alpha^2)=\alpha^2\big(\tfrac12\lambda_++o(\alpha^2)/\alpha^2\big)\gt0$ for small $\alpha\ne0$. In the same way $f(\x^\star+\alpha\u_-)\lt f(\x^\star)$. Both points lie in any given neighbourhood once $|\alpha|$ is small, so $\x^\star$ is a saddle. $\blacksquare$

The 2×2 shortcut. For $H=\begin{pmatrix}a&b\\b&d\end{pmatrix}$, $\det H=\lambda_1\lambda_2$ and $\operatorname{tr}H=\lambda_1+\lambda_2$. So: $\det H\lt0$ ⇒ saddle; $\det H\gt0$ and $a\gt0$ ⇒ minimum; $\det H\gt0$ and $a\lt0$ ⇒ maximum; $\det H=0$ ⇒ inconclusive.

Try it

This is the local picture near a critical point: $f\approx\tfrac12\lambda_1z_1^2+\tfrac12\lambda_2z_2^2$, where $z_1,z_2$ are coordinates along the eigenvectors $\u_1,\u_2$. Blue regions are above $f(\x^\star)$, orange below. Visit all four regimes. Then set $\lambda_1=1$, $\lambda_2=0$ and switch the higher-order term: the same Hessian gives a non-strict minimum, a strict minimum, or a saddle.

Classify the four critical points of $f(x,y)=x^3-3x+y^3-3y^2$ found in Chapter 3.1 (Tutorial 1, Problem 4).

  1. $\hess f(x,y)=\begin{pmatrix}6x&0\\0&6y-6\end{pmatrix}$.

    Differentiate $\grad f=(3x^2-3,\ 3y^2-6y)$ once more. No mixed terms, so the Hessian is diagonal and its eigenvalues are the diagonal entries.

  2. $(1,0)$: eigenvalues $6,-6$: indefinite, saddle. $\ (1,2)$: $6,6$: strict local minimum, $f=-6$.

    Read off the signs.

  3. $(-1,0)$: $-6,-6$: strict local maximum, $f=2$. $\ (-1,2)$: $-6,6$: saddle.

    Two saddles, one min, one max.

  4. The local minimum is not global: $f(0,y)=y^3-3y^2\to-\infty$ as $y\to-\infty$.

    Second-order conditions are local. A cubic is unbounded below, so there is no global minimizer at all.

Classify the critical points of $f(x_1,x_2)=2x_1^3+x_1x_2^2+5x_1^2+x_2^2$ (Tutorial 1, Problem 5: coupled cross-derivatives).

  1. $\grad f=\begin{pmatrix}6x_1^2+x_2^2+10x_1\\2x_2(x_1+1)\end{pmatrix}$.

    $x_1x_2^2$ contributes $x_2^2$ to the first partial and $2x_1x_2$ to the second.

  2. Second equation: $x_2=0$ or $x_1=-1$. If $x_2=0$: $2x_1(3x_1+5)=0$, so $x_1=0$ or $-\tfrac53$. If $x_1=-1$: $6+x_2^2-10=0$, so $x_2=\pm2$. Critical points: $(0,0)$, $(-\tfrac53,0)$, $(-1,2)$, $(-1,-2)$.

    Coupled equations: split on the factored one, then substitute each case into the other.

  3. $\hess f=\begin{pmatrix}12x_1+10&2x_2\\2x_2&2x_1+2\end{pmatrix}$. At $(0,0)$: $\operatorname{diag}(10,2)\succ0$: strict local min, $f=0$. At $(-\tfrac53,0)$: $\operatorname{diag}(-10,-\tfrac43)\prec0$: strict local max, $f=\tfrac{125}{27}\approx4.63$.

    With $x_2=0$ the off-diagonal entry vanishes, so the eigenvalues are the diagonal.

  4. At $(-1,\pm2)$: $\begin{pmatrix}-2&\pm4\\\pm4&0\end{pmatrix}$, $\det=-16\lt0$: saddles, $f=3$. (Eigenvalues $-1\pm\sqrt{17}\approx3.12,\ -5.12$.)

    A negative determinant means eigenvalues of opposite sign; no need to compute them.

Try it

The rings mark every critical point found by running Newton's method on $\grad f=\0$ from a grid of starts. Click near a ring (or press "Next critical point") to see the gradient, Hessian, eigenvalues, eigen-directions and verdict. Check both worked examples, then [LD]'s $x^3-x^2y+2y^2$ (Practice 2) and type your own, e.g. x^3 + y^3 - 3*x*y.

For $f(x,y)=x^4+y^4-4xy$, compute the eigenvalues of $\hess f(1,1)$. What kind of point is $(1,1)$?

$\hess f=\begin{pmatrix}12x^2&-4\\-4&12y^2\end{pmatrix}$. At $(1,1)$ the trace is 24 and the determinant is $144-16$.

$\hess f(1,1)=\begin{pmatrix}12&-4\\-4&12\end{pmatrix}$, eigenvalues $12\pm4$: 8 and 16. Both positive: a strict local minimum, $f(1,1)=-2$. (Same at $(-1,-1)$; the origin, with $\begin{pmatrix}0&-4\\-4&0\end{pmatrix}$ and eigenvalues $\pm4$, is a saddle.) In fact $-2$ is the global minimum: $4xy\le2(x^2+y^2)\le x^4+y^4+2$, because $x^4-2x^2+1=(x^2-1)^2\ge0$. So $f\ge-2$ everywhere.

([LD] §7.3, Example 2.) $f(x_1,x_2)=x_1^3-x_1^2x_2+2x_2^2$ has a critical point at $(6,9)$. With $x_2=9$ fixed, $x_1=6$ minimizes $f$; with $x_1=6$ fixed, $x_2=9$ minimizes $f$. Classify $(6,9)$.

$\hess f=\begin{pmatrix}6x_1-2x_2&-2x_1\\-2x_1&4\end{pmatrix}$. Evaluate at $(6,9)$ and look at the determinant.

$\hess f(6,9)=\begin{pmatrix}18&-12\\-12&4\end{pmatrix}$, $\det=72-144=-72\lt0$: indefinite, so a saddle (eigenvalues $11\pm\sqrt{193}\approx24.9,\ -2.9$). Being a minimum along each coordinate axis is not enough: the downhill eigen-direction is diagonal.

For $f(x,y)=x^2+kxy+y^2$, the origin is a strict local minimum when $|k|\lt2$ (Sylvester: $4-k^2\gt0$). What is the origin when $k=2$?

The Hessian $\begin{pmatrix}2&2\\2&2\end{pmatrix}$ is singular, so the test is silent. Look at $f$ directly: it's a perfect square.

$f=(x+y)^2\ge0=f(0,0)$, so the origin is a (global) minimizer. But $f=0$ along the whole line $y=-x$, so it is not strict. The Hessian has eigenvalues 0 and 4: SONC holds, SOSC is silent, and here the answer is a non-strict minimum.

For $f(x_1,x_2)=2x_1^3+x_1x_2^2+5x_1^2+x_2^2$, what is the value of $f$ at its local maximizer?

The local maximizer is $(-\tfrac53,0)$ (second worked example).

$f(-\tfrac53,0)=2\left(-\tfrac{125}{27}\right)+5\cdot\tfrac{25}{9}=-\tfrac{250}{27}+\tfrac{375}{27}=\tfrac{125}{27}\approx4.63$.

(Book, practice problem "A degenerate case".) At the origin, $f(x,y)=x^2+y^4$ has $\grad f=\0$. Which statement is correct?

Compute $\hess f(0,0)$. Then look at $f$ itself: can it ever be negative?

$\hess f(0,0)=\operatorname{diag}(2,0)$: positive semidefinite and singular, so SONC holds and SOSC does not apply. Directly, $f=x^2+y^4\ge0=f(0,0)$ with equality only at the origin: a strict global minimum, which the Hessian alone could not certify.

  • Follow the routine: gradient → solve $\grad f=\0$ (all cases) → Hessian at each point → eigenvalue signs → verdict
  • Use $\det$ and trace for $2\times2$ Hessians
  • Say which theorem you used: SOSC for min/max, "indefinite ⇒ saddle" for saddles
  • When the Hessian is singular, say "inconclusive" and examine $f$ directly
  • Calling a point a minimum because $H\succeq0$ (that's SONC: necessary only)
  • Calling a local minimum global without a separate argument
  • Deciding definiteness from the diagonal entries alone ([LD]'s $(6,9)$ has diagonal $18,4$ and is a saddle)
  • Mixing up the tools: Lagrange form for necessary conditions, asymptotic form for sufficient ones
  1. SONC: local min ⇒ $\grad f=\0$ and $\hess f\succeq0$ (proof: Lagrange form, then $t\to0$ using continuity of $\hess f$).
  2. SOSC: $\grad f=\0$ and $\hess f\succ0$ ⇒ strict local min (proof: asymptotic form plus Rayleigh $\p^\top H\p\ge\lambda_{\min}\norm{\p}^2$).
  3. Eigenvalue signs classify a critical point: all $+$ min, all $-$ max, mixed saddle, some zero inconclusive.

At a critical point, $\hess f(\x^\star)$ has eigenvalues $3$, $0$, $5$. What can you conclude?

$\x^\star$ is a strict local minimum
SOSC needs all eigenvalues strictly positive.
Nothing yet: SONC holds but the test is inconclusive
Semidefinite and singular. Along the zero-eigenvalue direction, higher-order terms decide.
$\x^\star$ is not a local minimum
It might be (think $x^2+y^4+z^2$). A zero eigenvalue doesn't rule a minimum out.

In the proof of SONC, why can't we simply evaluate the Hessian at $\x^\star$ from the start?

Because $\hess f(\x^\star)$ may not exist
$f\in C^2$, so it exists everywhere in $U$.
The exact (Lagrange) form puts the Hessian at an intermediate point; continuity of $\hess f$ is then used to pass to $\x^\star$
That's why the hypothesis is $C^2$ and why the limit $t\to0^+$ appears.
Because the Hessian is not symmetric
For $f\in C^2$ it is symmetric (Schwarz). That isn't the issue here.

Which step of the SOSC proof fails if $\hess f(\x^\star)$ is only positive semidefinite?

The Taylor expansion
The expansion holds for any $C^2$ function.
FONC
$\grad f(\x^\star)=\0$ is assumed in SOSC, not derived.
The bound $\tfrac12\lambda_{\min}\norm{\p}^2$ is then $0$, so it no longer dominates the $o(\norm{\p}^2)$ error
With $\lambda_{\min}=0$ the bracket is $0+o(1)$, whose sign is unknown.

A $2\times2$ Hessian at a critical point has $\det=-3$. The point is…

a local maximum
That needs two negative eigenvalues, whose product would be positive.
inconclusive
Inconclusive means $\det=0$.
a saddle point
$\lambda_1\lambda_2=-3\lt0$: one positive and one negative eigenvalue, so the Hessian is indefinite.

Saddles and the hard cases

Saddles are the typical critical points of non-convex functions; degenerate critical points (a zero eigenvalue) hide their nature from the Hessian; and Peano's surface shows that checking every straight line through a point is not enough.

These are favourite exam questions precisely because intuition fails, and they explain why algorithms can stall at saddles and how they escape.

A mountain pass: the lowest point of the ridge you're crossing, yet the highest point of the road you're walking. Whether it's "low" depends on which way you look.

Counting the downhill directions: the Morse index

A critical point is non-degenerate if $\hess f(\x^\star)$ is invertible (no zero eigenvalue), and degenerate otherwise.

The Morse index of a non-degenerate critical point is the number of negative eigenvalues of $\hess f(\x^\star)$, counted with multiplicity: the number of independent downhill directions.

If $\x^\star$ is a non-degenerate critical point of a smooth $f$ with Morse index $k$, there are smooth local coordinates $\z=(z_1,\dots,z_n)$ near $\x^\star$, with $\z(\x^\star)=\0$, in which $$f(\x)=f(\x^\star)-z_1^2-\dots-z_k^2+z_{k+1}^2+\dots+z_n^2.$$

So up to a smooth change of coordinates, every non-degenerate critical point is a quadratic with $k$ downhill and $n-k$ uphill directions: index $0$ is a minimum, index $n$ a maximum, anything in between a saddle. For example $f=x^2-y^2$ (Tutorial 1, Problem 6) has $\grad f=(2x,-2y)$, a single critical point at the origin with Hessian $\operatorname{diag}(2,-2)$: a saddle of index 1. The Morse lemma says nothing about degenerate points; those come later in this chapter.

Escaping a saddle along negative curvature

Take $f(x_1,x_2)=\tfrac12x_1^2-\tfrac12x_2^2+\tfrac14x_2^4$ (Tutorial 1, Problem 7). Then $\grad f=\big(x_1,\ x_2(x_2^2-1)\big)$ and $\hess f=\operatorname{diag}(1,\ 3x_2^2-1)$. The critical points are:

  • $(0,0)$: eigenvalues $1$ (eigenvector $(1,0)$) and $-1$ (eigenvector $(0,1)$): a saddle, Morse index 1.
  • $(0,\pm1)$: eigenvalues $1$ and $2$: strict local minima, index 0, with $f=-\tfrac14$. They are global: $f=\tfrac12x_1^2+\tfrac14(x_2^2-1)^2-\tfrac14\ge-\tfrac14$.

Gradient descent with step $\alpha$ updates the two coordinates separately: $$x_1\leftarrow(1-\alpha)x_1,\qquad x_2\leftarrow x_2\big(1+\alpha(1-x_2^2)\big).$$ Started exactly on the $x_1$-axis ($x_2=0$), $x_2$ stays 0 forever and, for $0\lt\alpha\lt2$, $x_1\to0$: gradient descent converges to the saddle. Nudge $x_2$ to any $\epsilon\ne0$: near the saddle $|x_2|$ is multiplied by about $1+\alpha$ each step, so it grows geometrically along the negative-curvature eigenvector until it reaches the minimum at $(0,\pm1)$ (for $\alpha\lt1$; larger steps overshoot it). One step from $(0,\epsilon)$ already lowers $f$: $f(0,\epsilon)=-\tfrac12\epsilon^2+\tfrac14\epsilon^4\lt0$ for $0\lt|\epsilon|\lt\sqrt2$. A random perturbation almost surely has a component along $\u_-$, which is the idea behind perturbed gradient descent.

Try it

Start gradient descent on the axis ("Start (1.2, 0)"): it slides into the saddle. Press "Nudge" and count the steps it takes to escape. Change $\alpha$ and watch the escape speed (about $\times(1+\alpha)$ per step). Then switch to pure Newton from $(0.5,0.3)$: it dives straight into the saddle. Drag the start above $x_2\approx0.58$ to see Newton behave.

Why plain Newton walks into saddles

Newton's method (covered properly in Part 10) jumps to the critical point of the local quadratic model: $\x_{k+1}=\x_k-[\hess f(\x_k)]^{-1}\grad f(\x_k)$. It is a root-finder for $\grad f=\0$, and it doesn't care whether that root is a minimum. Near the saddle above, $3x_2^2-1\lt0$ and the Newton update in $x_2$ is $x_2\leftarrow2x_2^3/(3x_2^2-1)$, which pulls $x_2$ towards 0 very fast.

On $f(x,y)=\tfrac12(x^2-y^2)$ (Tutorial 1, Problem 10): (a) take one pure Newton step from any $(x_0,y_0)$; (b) show that at $(1,2)$ the Newton direction goes uphill.

  1. $\grad f=(x,-y)^\top$, $\hess f=\operatorname{diag}(1,-1)$, so $[\hess f]^{-1}=\operatorname{diag}(1,-1)$.

    The Hessian is constant and invertible but indefinite.

  2. $\x_1=\begin{pmatrix}x_0\\y_0\end{pmatrix}-\begin{pmatrix}1&0\\0&-1\end{pmatrix}\begin{pmatrix}x_0\\-y_0\end{pmatrix}=\begin{pmatrix}0\\0\end{pmatrix}$.

    From every start, one step lands exactly on the saddle: the quadratic model is exact, and its only critical point is the saddle.

  3. At $(1,2)$: $\grad f=(1,-2)^\top$ and $\d_{\text{N}}=-[\hess f]^{-1}\grad f=-(1,2)^\top=(-1,-2)^\top$.

    The Newton direction.

  4. $\grad f^\top\d_{\text{N}}=(1)(-1)+(-2)(-2)=3\gt0$: an ascent direction. Indeed $f$ rises from $f(1,2)=-1.5$ to $f(0,0)=0$.

    $\grad f^\top\d_{\text{N}}=-\grad f^\top H^{-1}\grad f$ is negative when $H\succ0$; with an indefinite $H$ it can have either sign.

The standard remedies (Tutorial 1, Problem 10(c)) all put curvature-sign information back in: damping (solve $(\hess f+\tau I)\d=-\grad f$ with $\tau\gt-\lambda_{\min}$, which tends to steepest descent as $\tau\to\infty$); eigenvalue rectification (replace $H=U\Lambda U^\top$ by $U|\Lambda|U^\top$, which on this example is $I$ and turns the $y$-step around); and trust regions (minimize the model within $\norm{\d}\le\Delta$; with negative curvature the solution lies on the boundary, along the most negative eigenvector).

Degenerate critical points: when the Hessian says nothing

If $\hess f(\x^\star)$ has a zero eigenvalue, anything can happen. One-variable warm-up: $x^4$ (minimum), $-x^4$ (maximum), $x^3$ (neither), all with $f'(0)=f''(0)=0$. Two-variable versions, all with $\hess f(0,0)=\operatorname{diag}(2,0)$: $x^2+y^4$ is a strict minimum; $x^2-y^4$ and $x^2+y^3$ are saddles (look along the $y$-axis); $x^2$ alone has a whole line of minimizers.

The monkey saddle (Tutorial 1, Problem 8): find the critical points of $f(x,y)=x^3-3xy^2$ and classify them.

  1. $\grad f=(3x^2-3y^2,\ -6xy)$. The second entry is zero when $x=0$ or $y=0$; either way the first then forces the other coordinate to be 0. Only critical point: $(0,0)$.

    Split on the product, then substitute.

  2. $\hess f=\begin{pmatrix}6x&-6y\\-6y&-6x\end{pmatrix}$, so $\hess f(0,0)=0$: both eigenvalues are 0. SONC holds trivially, SOSC says nothing, and the Morse lemma doesn't apply.

    The zero matrix is both positive and negative semidefinite.

  3. Polar coordinates $x=r\cos\theta$, $y=r\sin\theta$: $f=r^3(\cos^3\theta-3\cos\theta\sin^2\theta)=r^3(4\cos^3\theta-3\cos\theta)=r^3\cos3\theta$.

    Use $\sin^2\theta=1-\cos^2\theta$, then the triple-angle identity $\cos3\theta=4\cos^3\theta-3\cos\theta$.

  4. For any $r\gt0$, $\cos3\theta$ takes both signs: six sectors of $60^\circ$, alternately $f\gt0$ and $f\lt0$. So every neighbourhood of the origin has points above and below $f(0,0)=0$: a saddle, revealed only by third-order terms.

    Three dips (two for the legs, one for the tail) give the name. A degenerate saddle can have more than two downhill "valleys".

Peano's surface: a minimum along every line, and still not a minimum

The book flags this as a favourite exam question. Let $$f(x,y)=(y-x^2)(y-2x^2)=y^2-3x^2y+2x^4.$$ The two parabolas $y=x^2$ and $y=2x^2$ are where $f=0$. Between them (in the thin sliver $x^2\lt y\lt2x^2$) one factor is positive and the other negative, so $f\lt0$; everywhere else $f\ge0$.

Show that the origin is a critical point where the Hessian test is inconclusive, that the origin is a strict local minimum along every straight line through it, and that it is nevertheless a saddle point.

  1. $\grad f=(-6xy+8x^3,\ 2y-3x^2)$, which is $\0$ at the origin. $\hess f=\begin{pmatrix}-6y+24x^2&-6x\\-6x&2\end{pmatrix}$, so $\hess f(0,0)=\operatorname{diag}(0,2)$: positive semidefinite and singular. SONC holds; SOSC is silent.

    Degenerate: the $x$-direction has zero curvature.

  2. Take any direction $\d=(d_1,d_2)\ne\0$ and $\phi(t)=f(td_1,td_2)=t^2d_2^2-3t^3d_1^2d_2+2t^4d_1^4$.

    Restrict $f$ to the line through the origin in direction $\d$.

  3. If $d_2\ne0$: $\phi(t)=t^2\big(d_2^2-3td_1^2d_2+2t^2d_1^4\big)$, and the bracket tends to $d_2^2\gt0$ as $t\to0$, so $\phi(t)\gt0=\phi(0)$ for small $t\ne0$. If $d_2=0$: $\phi(t)=2d_1^4t^4\gt0$ for $t\ne0$. Every line sees a strict minimum at $t=0$.

    The two cases cover every direction. The lowest-order nonzero term decides near $t=0$.

  4. But along the parabola $y=\tfrac32x^2$, which runs inside the sliver: $f\big(x,\tfrac32x^2\big)=\big(\tfrac12x^2\big)\big(-\tfrac12x^2\big)=-\tfrac14x^4\lt0$ for every $x\ne0$.

    Arbitrarily close to the origin there are points below $f(0,0)=0$; points above exist along any line. So the origin is a saddle by definition.

  5. Why the lines miss it: on the line in direction $\d$ (with $d_1,d_2\ne0$), $\phi(t)\gt0$ only for $0\lt|t|\lt|d_2|/(2d_1^2)$. That radius shrinks to 0 as the line approaches the $x$-axis, so no single ball around the origin works for all lines.

    A line $y=mx$ leaves the sliver, whose width shrinks like $x^2$, almost at once. "Minimum along every line" gives a radius for each line, but no common radius.

Try it

Orange shading is where $f\lt0$: the sliver between the parabolas. Rotate the purple line: its slice (the small plot) always has a strict minimum at $t=0$, and the thick orange segment shows where the line passes through the sliver. Bring the angle close to $0^\circ$ and watch the safe radius collapse. Then zoom in: the dashed curve $y=1.5x^2$ stays in the sliver at every scale, and its slice is always $-x^4/4$.

Careful: two different "Peano" examples appear in this course. Peano's surface $(y-x^2)(y-2x^2)$ (this chapter) is about optimality along lines. Peano's mixed-partials example $\frac{xy(x^2-y^2)}{x^2+y^2}$ (Tutorial 1, Problem 12; Part 0b) has $\frac{\partial^2f}{\partial y\partial x}(0,0)=-1\ne1=\frac{\partial^2f}{\partial x\partial y}(0,0)$, showing that Schwarz's theorem needs continuous second partials, i.e. why the second-order theory assumes $f\in C^2$.
Go deeper: Schwarz's theorem via difference boxes (Tutorial 1, Problem 11)

Let $\Delta(h,k)=f(h,k)-f(h,0)-f(0,k)+f(0,0)$. Write it as $\varphi(h)-\varphi(0)$ with $\varphi(x)=f(x,k)-f(x,0)$ and apply the mean value theorem twice (first in $x$, then in $y$) to get $\Delta/(hk)=\partial_{yx}f(\xi,\eta)$ for some $\xi\in(0,h)$, $\eta\in(0,k)$. Grouping the other way, $\psi(y)=f(h,y)-f(0,y)$, gives $\Delta/(hk)=\partial_{xy}f(\tilde\xi,\tilde\eta)$. Let $(h,k)\to(0,0)$: both points tend to the origin, and continuity of the mixed partials there forces $\partial_{yx}f(0,0)=\partial_{xy}f(0,0)$. In Peano's mixed-partials example, $f=\frac{r^2}{4}\sin4\theta$ in polar form and the mixed partial keeps oscillating with $\theta$ as $r\to0$, so continuity fails and so does the conclusion.

Go deeper: the higher-derivative test in one variable

If $f'(x^\star)=\dots=f^{(k-1)}(x^\star)=0$ and $f^{(k)}(x^\star)\ne0$ (with $f\in C^k$), Taylor gives $f(x^\star+h)-f(x^\star)=\frac{f^{(k)}(x^\star)}{k!}h^k+o(h^k)$. If $k$ is even, the sign of $f^{(k)}(x^\star)$ decides: positive gives a strict minimum, negative a strict maximum. If $k$ is odd, $h^k$ changes sign and there is no extremum. So $x^4$ ($k=4$, $24\gt0$) is a minimum and $x^3$ ($k=3$) is not. In several variables there is no such simple rule: Peano's surface passes every line test.

For Peano's surface $f(x,y)=(y-x^2)(y-2x^2)$, evaluate $f$ at the point of the parabola $y=1.5x^2$ with $x=0.2$.

$y=1.5(0.04)=0.06$. Or use the formula $f(x,1.5x^2)=-x^4/4$.

$(0.06-0.04)(0.06-0.08)=(0.02)(-0.02)=-0.0004=-(0.2)^4/4$. Negative, although the point is only about $0.21$ from the origin.

On Peano's surface, take the line through the origin in direction $\d=(1,1)/\sqrt2$. For $t\gt0$, $\phi(t)=f(t\d)$ is positive for $0\lt t\lt r$. Find $r$.

$\phi(t)=t^2\big(d_2^2-3td_1^2d_2+2t^2d_1^4\big)$ with $d_1=d_2=1/\sqrt2$. Find the smallest positive root of the bracket.

The bracket is $\tfrac12-\tfrac{3}{2\sqrt2}t+\tfrac12t^2=\tfrac12\big(t^2-\tfrac{3}{\sqrt2}t+1\big)$, with roots $t=\tfrac{1}{\sqrt2}$ and $t=\sqrt2$. So $r=|d_2|/(2d_1^2)=\tfrac{1}{\sqrt2}\approx0.7071$; for $0.707\lt t\lt1.414$ the line is inside the sliver and $\phi\lt0$.

Find the Morse index of the critical point at the origin of $f(x,y,z)=x^2+4xy+y^2-z^2$.

$\hess f=\begin{pmatrix}2&4&0\\4&2&0\\0&0&-2\end{pmatrix}$. The top-left $2\times2$ block has eigenvalues $2\pm4$.

The eigenvalues are $6$, $-2$ (from the block) and $-2$ (from $z$). None is zero, so the point is non-degenerate, and two are negative: Morse index 2, a saddle with two independent downhill directions, $(1,-1,0)$ and $(0,0,1)$.

Gradient descent with $\alpha=0.5$ on $f=\tfrac12x_1^2-\tfrac12x_2^2+\tfrac14x_2^4$ starts at $(0,\,0.01)$. After how many steps does $|x_2|$ first exceed $0.5$?

$x_2\leftarrow x_2(1+0.5(1-x_2^2))$: a factor just under $1.5$ while $x_2$ is small. Compare with $0.01\cdot1.5^k$.

The factor is at most $1.5$, so after 9 steps $x_2\le0.01\cdot1.5^9\approx0.384\lt0.5$. Iterating exactly: $0.0150,\ 0.0225,\ 0.0337,\ 0.0506,\ 0.0758,\ 0.1135,\ 0.1695,\ 0.2519,\ 0.3698,\ 0.5294$. It first exceeds 0.5 at step 10. Escape is geometric along the negative-curvature direction.

For $f(x,y)=\tfrac12(x^2-y^2)$ at $(1,2)$, compute $\grad f^\top\d_{\text N}$ where $\d_{\text N}=-[\hess f]^{-1}\grad f$ is the Newton direction.

$\grad f(1,2)=(1,-2)$ and $[\hess f]^{-1}=\operatorname{diag}(1,-1)$.

$\d_{\text N}=-(1,2)=(-1,-2)$, so $\grad f^\top\d_{\text N}=-1+4=3\gt0$: the Newton direction is an ascent direction.

  • Check for zero eigenvalues before quoting the Morse lemma or SOSC
  • For degenerate points, look at $f$ itself: factor it, use polar coordinates, or restrict to cleverly chosen curves
  • To prove "saddle", exhibit points above and below arbitrarily close to $\x^\star$
  • Remember a random perturbation lets gradient descent escape a strict saddle
  • Concluding "minimum" because every straight-line slice has a minimum (Peano)
  • Concluding "minimum" because the point minimizes along each coordinate axis ([LD]'s $(6,9)$)
  • Trusting pure Newton on a non-convex function: it converges to any critical point, saddles included
  • Confusing Peano's surface with Peano's mixed-partials example
  1. Morse index = number of negative Hessian eigenvalues; at non-degenerate points the Morse lemma makes $f$ locally a pure quadratic with that many downhill directions.
  2. At degenerate points the Hessian is silent: $x^2\pm y^4$, the monkey saddle $r^3\cos3\theta$ and Peano's surface must be classified directly.
  3. Gradient descent escapes a strict saddle along the negative-curvature eigenvector unless started exactly on the stable set; pure Newton is attracted to saddles.

A non-degenerate critical point in $\R^4$ has Hessian eigenvalues $-2,\,1,\,3,\,-0.5$. Its Morse index and type are…

index 2, local maximum
A maximum needs index $n=4$: all eigenvalues negative.
index 2, saddle
Two negative eigenvalues; since $0\lt2\lt4$, it is a saddle.
index 0, local minimum
The index counts the negative eigenvalues.

Why doesn't "the origin is a strict local minimum along every line" make it a local minimum of Peano's surface?

Because the Hessian at the origin is singular
Singular Hessians don't stop $x^2+y^4$ from having a minimum. The issue is about the radius on each line.
Each line has its own safe radius, and these radii shrink to 0, so no single ball works
Lines near the $x$-axis enter the negative sliver ever closer to the origin, and the parabola $y=1.5x^2$ stays inside it.
Because one special line has $\phi(t)\lt0$
No line does: every line's slice has a strict minimum at $t=0$. Re-read step 3 of the worked example.

Gradient descent on $\tfrac12x_1^2-\tfrac12x_2^2+\tfrac14x_2^4$ with $\alpha=0.3$ starts at $(1,0)$. It…

escapes to $(0,1)$
To escape it needs a nonzero $x_2$-component. The update keeps $x_2=0$ forever.
converges to the saddle $(0,0)$
$x_1\leftarrow0.7x_1\to0$ while $x_2$ stays exactly 0: the start is on the saddle's stable set.
diverges
$|1-\alpha|=0.7\lt1$, so the $x_1$-iterates shrink.

For the monkey saddle $x^3-3xy^2$, how many sectors around the origin have $f\lt0$?

1
$f=r^3\cos3\theta$. How many times is $\cos3\theta$ negative as $\theta$ goes round once?
2
That's an ordinary saddle like $x^2-y^2$. Here the angle is multiplied by 3.
3
$\cos3\theta$ completes three periods: three negative and three positive $60^\circ$ sectors.

Which form of Taylor's theorem is the standard tool for proving SOSC?

The integral form
That one is used for global inequalities like the descent lemma (Part 6).
The Lagrange form
It's used for necessary conditions (SONC). For a sufficient condition we need the Hessian at $\x^\star$ itself.
The asymptotic form with $o(\norm{\p}^2)$
It puts $\hess f(\x^\star)$ in the quadratic term, and the Rayleigh bound then beats the error for small $\p$.

$f(x,y)=x^2+y^2-x^2y$ has a critical point at the origin. It is…

a strict local minimum but not a global one
$\hess f(0,0)=2I\succ0$, so SOSC gives a strict local min. But $f(x,2)=4-x^2\to-\infty$.
a strict global minimum
SOSC is local. Try $y=2$ and large $x$.
a saddle
Compute $\hess f(0,0)$: the cubic term contributes nothing at the origin.

Gradient descent started exactly at $x_0=0$ on $x^4-2x^2$ never moves. Which statement explains why it is not at a minimum?

FONC fails at 0
$f'(0)=0$, so FONC holds. That's why the method doesn't move.
SONC fails at 0: $f''(0)=-4\lt0$
A local minimizer needs $f''\ge0$. In fact $f''(0)\lt0$ makes 0 a strict local maximum.
0 is a degenerate critical point
$f''(0)=-4\ne0$: non-degenerate.

FONC for the fit-the-line loss $L(\z)=\tfrac12\z^\top A\z-\b^\top\z+38$ is the linear system…

$A\z=\0$
You've dropped the linear term: $\grad L=A\z-\b$.
$A\z=\b$, giving $\z^\star=(1.5,\,1/3)$
The normal equations. Since $A\succ0$ (eigenvalues ≈0.72 and 33.3), SOSC makes it a strict local minimizer.
$2A\z=\b$
The $\tfrac12$ cancels the 2 from differentiating the quadratic form.

Pure Newton's method, run on a function with a saddle and a minimum…

always converges to a minimum, since it uses curvature
It uses curvature to find a zero of $\grad f$, not to choose between kinds of critical point.
can converge to the saddle, and its direction can even point uphill
On $\tfrac12(x^2-y^2)$ one step lands on the saddle from anywhere, and at $(1,2)$ $\grad f^\top\d_{\text N}=3\gt0$.
can't be applied, because the Hessian is indefinite
An indefinite Hessian can still be invertible, so the step is defined. That's what makes it dangerous.

The book's "Interlude: the convexity toolkit", which every lecture from Lecture 6 onwards leans on. You'll learn what convex sets and functions are, three equivalent ways to test convexity (with the proofs the exam asks for), why a local minimizer of a convex function is automatically global, and the hierarchy convex → strictly convex → strongly convex, ending with the function classes $\mathcal F_L^{1,1}$ and $\mathcal S_{\mu,L}^{1,1}$ and the condition number $L/\mu$ that drives Part 5.

You need: Part 0b (gradient, Hessian, eigenvalues and definiteness, Taylor's theorem), Part 1 (local vs global minimizers), and the first-order condition $\grad f(\x^\star)=\0$ and the Weierstrass theorem (Parts 2–3).

Convex sets and convex functions

A set is convex if the straight segment between any two of its points stays inside it; a function is convex if the straight chord between any two points on its graph stays on or above the graph.

Part 1 quoted Rockafellar: the real watershed in optimization is convexity versus non-convexity. Convex problems are the ones where the local information an algorithm sees is enough to find the global answer. This chapter gives you the definitions you'll use in every proof from here on.

A bowl versus an egg box. Drop a marble anywhere in a bowl and it ends up at the one lowest point. Drop it in an egg box and it gets stuck in whichever cup it lands in.

Convex sets: no dents, no holes

Take two points $\x$ and $\y$. The points on the straight segment between them are exactly the convex combinations $\theta\x+(1-\theta)\y$ with $\theta\in[0,1]$: at $\theta=1$ you're at $\x$, at $\theta=0$ you're at $\y$, and $\theta=\tfrac12$ is the midpoint. (In one variable, $\theta\cdot3+(1-\theta)\cdot7$ sweeps from 7 down to 3 as $\theta$ goes from 0 to 1.)

A set $C\subseteq\R^n$ is convex if for all $\x,\y\in C$ and all $\theta\in[0,1]$, $$\theta\x+(1-\theta)\y\in C.$$ In words: whenever two points are in $C$, the whole segment joining them is in $C$.

Examples. All of $\R^n$; a ball $\{\x:\norm{\x-\mathbf c}\le r\}$; a half-space $\{\x:\a^\top\x\le\beta\}$; a line; a single point; the empty set. Non-examples. A ring (annulus): the segment between two opposite points passes through the hole. Two separate discs: the segment between them crosses the gap. A crescent or a peanut shape: the segment cuts across the dent.

Intersections stay convex. If $C_1$ and $C_2$ are convex and $\x,\y$ lie in both, the segment lies in $C_1$ (since $C_1$ is convex) and in $C_2$, so it lies in $C_1\cap C_2$. This is why the feasible region of Part 1's diet problem, an intersection of half-planes, is a convex polygon.

Try it

Pick a set and drag the two endpoints. For the disc and triangle, try hard to make the segment leave the set: you can't. For the annulus, peanut and two discs, find a pair whose segment escapes. One bad pair is enough to prove a set is not convex; "Search for a bad pair" automates the hunt.

Convex functions: the chord lies above the graph

Now take a function $f$ and two inputs $\x,\y$. Draw the straight line (the chord) from the point $(\x,f(\x))$ on the graph to the point $(\y,f(\y))$. At the input $\theta\x+(1-\theta)\y$, the chord has height $\theta f(\x)+(1-\theta)f(\y)$, the same mix of the two end heights. Convexity says the graph never rises above the chord.

Let $C\subseteq\R^n$ be a convex set. A function $f:C\to\R$ is convex if for all $\x,\y\in C$ and all $\theta\in[0,1]$, $$f(\theta\x+(1-\theta)\y)\ \le\ \theta f(\x)+(1-\theta)f(\y).$$ $f$ is concave if $-f$ is convex.

Equivalently, the epigraph $\operatorname{epi}f=\{(\x,t)\in C\times\R:\ t\ge f(\x)\}$, everything on or above the graph, is a convex set in $\R^{n+1}$.

Symbols
$\theta$
the mixing weight, between 0 and 1
$\theta\x+(1-\theta)\y$
a point on the segment from $\y$ ($\theta=0$) to $\x$ ($\theta=1$)
$\theta f(\x)+(1-\theta)f(\y)$
the chord's height above that point
$\operatorname{epi}f$
"epigraph": the region on or above the graph

The domain must be convex, otherwise the point $\theta\x+(1-\theta)\y$ might not even be in it. This is why the definition starts with "let $C$ be a convex set".

Why the epigraph? If the graph had a dent, the region above it would have a dent too, and a segment between two points of the region could pass under the graph. The epigraph turns "convex function" into "convex set", so pictures of sets apply to functions.

Try it

Drag the two blue points along the graph. For $x^2$, $e^x$ and $|x|$ the chord never dips below the graph. For $x^4-2x^2$ and $\sin x$, find placements where the graph pokes above the chord (shaded orange). For $|x|$, put both points on the same side of 0: the chord lies on the graph. That's allowed: the definition says $\le$. Then type your own function, e.g. x^4 + x or log(1+exp(x)).

Prove from the definition that $f(x)=x^2$ is convex on $\R$, and find exactly how far below the chord the graph sits.

  1. Compute the gap between chord and graph: $G=\theta x^2+(1-\theta)y^2-\big(\theta x+(1-\theta)y\big)^2$.

    Convexity is the statement $G\ge0$ for every $x,y$ and $\theta\in[0,1]$, so we just compute $G$.

  2. Expand the square: $(\theta x+(1-\theta)y)^2=\theta^2x^2+2\theta(1-\theta)xy+(1-\theta)^2y^2$. Then $\theta x^2-\theta^2x^2=\theta(1-\theta)x^2$ and $(1-\theta)y^2-(1-\theta)^2y^2=\theta(1-\theta)y^2$.

    Group the $x^2$ terms and the $y^2$ terms; each leaves a factor $\theta(1-\theta)$.

  3. So $G=\theta(1-\theta)\,(x^2-2xy+y^2)=\theta(1-\theta)(x-y)^2\ge0$.

    Both factors are nonnegative when $\theta\in[0,1]$. Hence $x^2$ is convex.

  4. Moreover $G>0$ whenever $x\ne y$ and $0\lt\theta\lt1$: the graph is strictly below the chord. And the gap is at least $\theta(1-\theta)(x-y)^2$, a fixed multiple of the squared distance.

    These two extra facts are exactly "strictly convex" and "strongly convex", the hierarchy of Chapter 4.3. Keep this computation in mind.

A catalogue of convex functions

  • Affine functions $f(\x)=\a^\top\x+\beta$: the definition holds with equality, so they are convex and concave. They're the boundary case.
  • Every norm, e.g. $\norm{\x}_2$, $\norm{\x}_1$, $|x|$: by the triangle inequality and $\norm{c\v}=|c|\norm{\v}$, $\ \norm{\theta\x+(1-\theta)\y}\le\theta\norm{\x}+(1-\theta)\norm{\y}$.
  • Convex quadratics $\tfrac12\x^\top A\x-\b^\top\x+c$ with $A\succeq0$ (proved in Chapter 4.2), including the fit-the-line loss $L(w,c)$ and least squares $\tfrac12\norm{A\x-\b}^2$.
  • One-variable classics: $e^x$; $x^p$ on $[0,\infty)$ for $p\ge1$; $-\log x$ and $x\log x$ on $(0,\infty)$.
  • Log-sum-exp $\log\sum_ie^{x_i}$, the smooth version of $\max_i x_i$ used in machine learning (proved in Chapter 4.2).

Not convex: $x^4-2x^2$ (two valleys), $\sin x$, $x^3$ (concave for $x\lt0$), $\sqrt{|x|}$, and $x_1x_2$ (a saddle).

Let $f,g$ be convex on a convex set $C$.

  1. Nonnegative combinations: $af+bg$ is convex for $a,b\ge0$ ([LD] §7.4 Props. 1–2).
  2. Affine change of variable: $h(\x)=f(M\x+\mathbf c)$ is convex (for any matrix $M$ and vector $\mathbf c$).
  3. Pointwise maximum: $\max\{f(\x),g(\x)\}$ is convex.
  4. Sublevel sets are convex: $\{\x\in C: f(\x)\le c\}$ is a convex set for every $c$ ([LD] §7.4 Prop. 3). So constraints $f_1(\x)\le c_1,\dots,f_m(\x)\le c_m$ with convex $f_i$ define a convex feasible set.

Why (2): $M(\theta\x+(1-\theta)\y)+\mathbf c=\theta(M\x+\mathbf c)+(1-\theta)(M\y+\mathbf c)$, so a segment in $\x$ maps to a segment in $M\x+\mathbf c$, and $f$'s chord inequality applies there. Why (3): $f(\theta\x+(1-\theta)\y)\le\theta f(\x)+(1-\theta)f(\y)\le\theta\max\{f,g\}(\x)+(1-\theta)\max\{f,g\}(\y)$, and the same for $g$; so the larger of the two is also below the right-hand side. Why (4): if $f(\x)\le c$ and $f(\y)\le c$ then $f(\theta\x+(1-\theta)\y)\le\theta c+(1-\theta)c=c$.

Careful: the converse of (4) is false. A function can have convex sublevel sets without being convex: $\sqrt{|x|}$ has sublevel sets $[-c^2,c^2]$, all intervals, yet it is concave on $(0,\infty)$. Products and minimums of convex functions need not be convex either: $x$ and $x^2$ are convex but $x\cdot x^2=x^3$ is not; $\min\{(x-1)^2,(x+1)^2\}$ has two valleys.
Go deeper: the epigraph equivalence, and composition rules

Convex $f\iff$ convex epigraph. ($\Rightarrow$) Take $(\x,s),(\y,t)$ in $\operatorname{epi}f$, so $s\ge f(\x)$, $t\ge f(\y)$. Then $\theta s+(1-\theta)t\ge\theta f(\x)+(1-\theta)f(\y)\ge f(\theta\x+(1-\theta)\y)$, so the mixed point is in the epigraph. ($\Leftarrow$) The points $(\x,f(\x))$ and $(\y,f(\y))$ are in the epigraph; if it is convex, so is their mix $(\theta\x+(1-\theta)\y,\ \theta f(\x)+(1-\theta)f(\y))$, which says exactly that the chord height is $\ge f$ at the mixed point.

Composition. If $g$ is convex and $\phi:\R\to\R$ is convex and nondecreasing, then $\phi(g(\x))$ is convex: $\phi(g(\theta\x+(1-\theta)\y))\le\phi(\theta g(\x)+(1-\theta)g(\y))\le\theta\phi(g(\x))+(1-\theta)\phi(g(\y))$, using "nondecreasing" for the first step and convexity of $\phi$ for the second. Example: $e^{\norm{\x}^2}$ is convex. Without "nondecreasing" it fails: $\phi(u)=-u$ is convex (it is linear) but decreasing, and $\phi(g)=-g$ is concave.

Sets. The set of minimizers of a convex function is the sublevel set at height $\min f$, so it is convex ([LD] §7.5 Thm. 1). A convex function can have one minimizer, none, or a whole segment of them, but never two isolated ones.

For $f(x)=e^x$, $x=0$, $y=2$ and $\theta=\tfrac12$, compute the gap "chord height minus function value" at the midpoint, $\tfrac12f(0)+\tfrac12f(2)-f(1)$.

The chord height is $\tfrac12(1+e^2)$; the function value at the midpoint $x=1$ is $e$.

$\tfrac12(1+e^2)-e=\tfrac12(1+7.3891)-2.7183=4.1945-2.7183\approx1.4762$. Positive, as convexity demands. (In fact it equals $\tfrac12(e-1)^2$.)

Is the set $\{\x\in\R^2:\ x_1^2+x_2^2\ge1\}$ (everything outside the open unit disc) convex?

Look for two points in the set whose midpoint is not. Try points on opposite sides of the origin.

$\x=(1,0)$ and $\y=(-1,0)$ are in the set, but their midpoint $(0,0)$ has $0^2+0^2=0\lt1$. One bad pair proves it: not convex. It is a superlevel set $\{\norm{\x}^2\ge1\}$ of a convex function, and superlevel sets of convex functions need not be convex.

Is $f(\x)=\max\{x_1^2,\ |x_2|,\ 3-x_1+2x_2\}$ convex on $\R^2$?

Check each piece separately, then use one of the preservation rules.

$x_1^2$ is convex (a convex function of $x_1$, i.e. composed with the linear map $\x\mapsto x_1$); $|x_2|$ is a norm of a linear map of $\x$, so convex; $3-x_1+2x_2$ is affine. The pointwise maximum of convex functions is convex. So $f$ is convex, even though it has kinks.

$f$ and $g$ are convex functions of one variable. Which of these is not guaranteed to be convex?

Three of them match a rule in the box. For the fourth, test $f(x)=x$ and $g(x)=x^2$.

$f+2g$ (nonnegative combination), $\max(f,g)$ (pointwise max) and $f(3x-1)$ (affine change of variable) are always convex. The product is not: $f(x)=x$ and $g(x)=x^2$ are convex but $fg=x^3$ has $(x^3)''=6x\lt0$ for $x\lt0$, and the chord from $-2$ to $0$ lies below the graph.

  • To prove a set or function is not convex, exhibit one explicit bad pair (and $\theta=\tfrac12$ usually suffices)
  • To prove convexity, build the function from known convex pieces with the preservation rules
  • Check that the domain is a convex set before talking about a convex function on it
  • Thinking "convex" means "curved upward somewhere"; it must hold for every pair of points
  • Assuming products, minimums or differences of convex functions are convex
  • Concluding $f$ is convex just because its sublevel sets are convex
  1. A set is convex if it contains the segment between any two of its points; intersections of convex sets are convex.
  2. $f$ is convex if $f(\theta\x+(1-\theta)\y)\le\theta f(\x)+(1-\theta)f(\y)$: every chord lies on or above the graph, equivalently the epigraph is convex.
  3. Nonnegative sums, affine substitutions and pointwise maxima preserve convexity; sublevel sets of convex functions are convex.

Which of these sets is not convex?

$\{\x\in\R^2: x_1+x_2\le1,\ x_1\ge0,\ x_2\ge0\}$
Each condition is a half-plane, and intersections of convex sets are convex.
$\{\x\in\R^2: x_1^2+4x_2^2\le4\}$
This is a sublevel set of a convex quadratic. What does the box say about those?
$\{\x\in\R^2: x_1x_2\ge1\}$
It has two branches, one in each of the first and third quadrants: $(1,1)$ and $(-1,-1)$ are in it, but their midpoint $(0,0)$ is not.

$f(x)=|x|$ satisfies the chord inequality with equality for $x=1$, $y=3$. This shows that…

$|x|$ is not convex
The definition uses $\le$; equality is allowed.
$|x|$ is convex but not strictly convex
Right: on a piece where $|x|$ is linear, the chord lies on the graph. Chapter 4.3 names this.
$|x|$ is concave
Concave would need $\ge$ for all pairs. Check $x=-1$, $y=1$.

Why must the domain of a convex function be a convex set?

Because convex functions are always defined on all of $\R^n$
$-\log x$ is convex on $(0,\infty)$ only.
Because otherwise $f$ would not be continuous
Continuity is not the issue here. Look at the left-hand side of the inequality.
So that $f(\theta\x+(1-\theta)\y)$ makes sense: the mixed point must be in the domain
The definition evaluates $f$ on the segment, so the segment must lie in the domain.

A feasible set is $\{\x: f_1(\x)\le0,\ f_2(\x)\le0\}$ with $f_1,f_2$ convex. It is…

always convex
Each constraint gives a convex sublevel set, and their intersection is convex.
convex only if $f_1,f_2$ are linear
Linearity is more than you need. Which property of sublevel sets applies?
convex only if it is bounded
Unbounded sets like half-planes are convex too.

Three equivalent tests for convexity

For a smooth function, "the chord lies above the graph" is equivalent to "every tangent lies below the graph", which is equivalent to "the Hessian is positive semidefinite everywhere".

The exam routinely asks you to prove these equivalences and to pick the right test: the Hessian test to check a given function, the tangent test to prove general theorems.

Three views of the same bowl: from above (chords sag less than the rim), from the side (every tangent plank lies under it) and from inside (it curves up in every direction).

First order: tangents are global under-estimators

Recall from Part 0b that $f(\x)+\grad f(\x)^\top(\y-\x)$ is the first-order Taylor approximation of $f(\y)$ around $\x$: the tangent line (or tangent plane). For a general function it is only good near $\x$. For a convex function something much stronger holds: it is a lower bound everywhere.

Let $C$ be open and convex and $f\in C^1(C)$. Then $f$ is convex if and only if $$f(\y)\ \ge\ f(\x)+\grad f(\x)^\top(\y-\x)\qquad\text{for all }\x,\y\in C.$$ Equivalently, the gradient is monotone: $$\big(\grad f(\y)-\grad f(\x)\big)^\top(\y-\x)\ \ge\ 0\qquad\text{for all }\x,\y\in C.$$

Let $C$ be open and convex and $f\in C^2(C)$. Then $f$ is convex if and only if $\hess f(\x)\succeq0$ (positive semidefinite) for every $\x\in C$.

In one variable these read: $f(y)\ge f(x)+f'(x)(y-x)$; $f'$ is nondecreasing ($(f'(y)-f'(x))(y-x)\ge0$ says $f'$ moves in the same direction as $x$); and $f''\ge0$. A function is convex exactly when its slope never decreases.

The payoff of the tangent inequality. If $\grad f(\x^\star)=\0$, it reads $f(\y)\ge f(\x^\star)$ for every $\y$. A flat tangent at a point of a convex function is a floor under the whole graph. Chapter 4.3 builds on this.

Try it

Drag the purple point: its tangent line moves with it. For a convex function the tangent never cuts above the graph, wherever you put it. For $x^4-2x^2$ and $\sin x$, find tangents that the graph dips below, and notice they sit where the strip at the bottom turns orange ($f''\lt0$). Switch to "Both" to see chords and tangents together.

The proofs (examinable)

Tutorial 1 Problems 14 and 16 ask for these. Each proof is short once you know the one trick it uses; learn the trick, not the algebra.

Prove: chord definition $\iff$ tangent inequality (for $f\in C^1$ on an open convex $C$).

  1. ($\Rightarrow$) Fix $\x,\y$ and $\theta\in(0,1]$. Rewrite the mixed point as $\x+\theta(\y-\x)=(1-\theta)\x+\theta\y$. The chord inequality gives $f(\x+\theta(\y-\x))\le(1-\theta)f(\x)+\theta f(\y)=f(\x)+\theta\big(f(\y)-f(\x)\big)$.

    Trick: write the mixed point as "start at $\x$, walk a fraction $\theta$ towards $\y$".

  2. Subtract $f(\x)$ and divide by $\theta\gt0$: $\dfrac{f(\x+\theta(\y-\x))-f(\x)}{\theta}\le f(\y)-f(\x)$.

    The left side is a difference quotient, the kind whose limit is a derivative.

  3. Let $\theta\to0^+$. The left side tends to the directional derivative $\grad f(\x)^\top(\y-\x)$ (Part 0b). So $\grad f(\x)^\top(\y-\x)\le f(\y)-f(\x)$, which is the tangent inequality.

    Inequalities survive limits (a $\le$ that holds for every $\theta$ holds in the limit).

  4. ($\Leftarrow$) Fix $\x,\y$, $\theta\in[0,1]$, and let $\z=\theta\x+(1-\theta)\y$. Apply the tangent inequality at $\z$, twice: $f(\x)\ge f(\z)+\grad f(\z)^\top(\x-\z)$ and $f(\y)\ge f(\z)+\grad f(\z)^\top(\y-\z)$.

    Trick: put the tangent at the mixed point and look outwards at both ends.

  5. Multiply the first by $\theta$, the second by $1-\theta$, and add: $\theta f(\x)+(1-\theta)f(\y)\ge f(\z)+\grad f(\z)^\top\big(\theta\x+(1-\theta)\y-\z\big)=f(\z)$.

    The bracket is $\z-\z=\0$, so the gradient term vanishes. What's left is the chord inequality. ([LD] §7.4 Prop. 4.)

Prove: tangent inequality $\iff$ $\hess f\succeq0$ everywhere (for $f\in C^2$ on an open convex $C$).

  1. ($\Rightarrow$) Fix $\x\in C$ and any direction $\v$. For small $t\gt0$, $\x+t\v\in C$ (because $C$ is open). Taylor with Lagrange remainder: $f(\x+t\v)=f(\x)+t\,\grad f(\x)^\top\v+\tfrac{t^2}2\v^\top\hess f(\x+\xi t\v)\v$ for some $\xi\in(0,1)$.

    The Lagrange form proves necessary conditions (Part 0b's rule of thumb), and "convex $\Rightarrow$ PSD" is a necessary condition.

  2. The tangent inequality with $\y=\x+t\v$ says $f(\x+t\v)\ge f(\x)+t\,\grad f(\x)^\top\v$. Subtracting, $\tfrac{t^2}2\v^\top\hess f(\x+\xi t\v)\v\ge0$.

    The tangent inequality says exactly that everything beyond the linear term is nonnegative.

  3. Divide by $t^2/2$ and let $t\to0^+$. Since $\hess f$ is continuous, $\v^\top\hess f(\x)\v\ge0$. As $\v$ was arbitrary, $\hess f(\x)\succeq0$.

    This is where $f\in C^2$ (continuous second derivatives) is used.

  4. ($\Leftarrow$) For any $\x,\y\in C$, Taylor gives $f(\y)=f(\x)+\grad f(\x)^\top(\y-\x)+\tfrac12(\y-\x)^\top\hess f(\boldsymbol\xi)(\y-\x)$ with $\boldsymbol\xi$ on the segment (in $C$, since $C$ is convex). The last term is $\ge0$ because $\hess f(\boldsymbol\xi)\succeq0$. Drop it: $f(\y)\ge f(\x)+\grad f(\x)^\top(\y-\x)$.

    Here PSD must hold at every point, because we don't know where $\boldsymbol\xi$ is. PSD at one point is not enough.

Prove: tangent inequality $\iff$ monotone gradient (Tutorial 1 Problem 16).

  1. ($\Rightarrow$) Write the tangent inequality twice, swapping roles: $f(\y)\ge f(\x)+\grad f(\x)^\top(\y-\x)$ and $f(\x)\ge f(\y)+\grad f(\y)^\top(\x-\y)$.

    Trick: use the inequality from both ends.

  2. Add them and cancel $f(\x)+f(\y)$: $0\ge\grad f(\x)^\top(\y-\x)-\grad f(\y)^\top(\y-\x)$, i.e. $\big(\grad f(\y)-\grad f(\x)\big)^\top(\y-\x)\ge0$.

    That is monotonicity.

  3. ($\Leftarrow$) Let $g(t)=f(\x+t(\y-\x))$ on $[0,1]$. By the chain rule $g'(t)=\grad f(\x+t(\y-\x))^\top(\y-\x)$. For $t\gt0$ apply monotonicity to the pair $\z=\x+t(\y-\x)$ and $\x$: $\big(\grad f(\z)-\grad f(\x)\big)^\top t(\y-\x)\ge0$, so (dividing by $t$) $g'(t)\ge g'(0)$.

    Trick: reduce to one variable along the segment, where monotone gradient means "the slope only grows".

  4. By the fundamental theorem of calculus, $f(\y)-f(\x)=g(1)-g(0)=\int_0^1g'(t)\,dt\ge\int_0^1g'(0)\,dt=\grad f(\x)^\top(\y-\x)$.

    This is the integral form of Taylor (Part 0b), the form for global inequalities.

Verify that a specific function is convex: compute $\hess f$ and show it is PSD everywhere (eigenvalues, Sylvester, or $\v^\top\hess f\,\v\ge0$ directly). Prove a general theorem about convex functions: use the tangent inequality. Disprove convexity: find one bad chord, or one point where $\hess f$ has a negative eigenvalue.

Applying the Hessian test

Quadratics. $f(\x)=\tfrac12\x^\top A\x-\b^\top\x+c$ has $\hess f=A$ at every point, so $f$ is convex $\iff A\succeq0$. Least squares $\tfrac12\norm{A\x-\b}^2$ has Hessian $A^\top A$, and $\v^\top A^\top A\v=\norm{A\v}^2\ge0$: always convex. The fit-the-line loss has Hessian $\begin{pmatrix}28&12\\12&6\end{pmatrix}$ with eigenvalues $\approx0.721$ and $33.28$, both positive: convex.

Prove that log-sum-exp, $f(\x)=\log\sum_{i=1}^ne^{x_i}$, is convex on $\R^n$ (Tutorial 1 Problem 15).

  1. Let $S=\sum_ke^{x_k}$ and $p_i=e^{x_i}/S$. Then $\partial f/\partial x_i=e^{x_i}/S=p_i$. Note $p_i\gt0$ and $\sum_ip_i=1$: the $p_i$ are the "softmax" probabilities.

    Chain rule: the derivative of $\log S$ is $S'/S$.

  2. Differentiate $p_i=e^{x_i}/S$ by the quotient rule: $\partial p_i/\partial x_i=p_i-p_i^2$ and $\partial p_i/\partial x_j=-p_ip_j$ for $j\ne i$. So $\hess f=\operatorname{diag}(\p)-\p\p^\top$.

    The diagonal gets the extra term because $e^{x_i}$ itself depends on $x_i$.

  3. For any $\v$: $\v^\top\hess f\,\v=\sum_ip_iv_i^2-\big(\sum_ip_iv_i\big)^2$.

    $\v^\top\operatorname{diag}(\p)\v=\sum p_iv_i^2$ and $\v^\top\p\p^\top\v=(\p^\top\v)^2$.

  4. Cauchy–Schwarz with $a_i=\sqrt{p_i}\,v_i$, $b_i=\sqrt{p_i}$: $\big(\sum_ip_iv_i\big)^2=\big(\sum a_ib_i\big)^2\le\big(\sum a_i^2\big)\big(\sum b_i^2\big)=\sum_ip_iv_i^2\cdot1$. Hence $\v^\top\hess f\,\v\ge0$, so $\hess f\succeq0$ everywhere, and $f$ is convex.

    The split $p_iv_i=\sqrt{p_i}v_i\cdot\sqrt{p_i}$ is chosen so that $\sum b_i^2=\sum p_i=1$. (In probability language: the quantity is a variance, which is never negative.)

Careful: "$\hess f(\x^\star)\succ0$ at one point" (the second-order sufficient condition of Part 3) is a local statement: $f$ curves up near $\x^\star$. Convexity needs PSD at every point. And $f''\gt0$ everywhere implies strictly convex, but not conversely: $x^4$ is strictly convex with $f''(0)=0$.

For which values of $a$ is $f(x)=x^4+a\,x^2$ convex on $\R$? Enter the smallest such $a$.

$f''(x)=12x^2+2a$. Where is it smallest?

$f''(x)=12x^2+2a\ge2a$, with equality at $x=0$. So $f''\ge0$ everywhere iff $a\ge0$. The smallest is $a=0$ (where $f=x^4$, still convex). For $a\lt0$, $f''(0)=2a\lt0$: not convex, and indeed $x^4+ax^2$ then has two valleys.

$f(x_1,x_2)=x_1^2x_2^2$ is convex in $x_1$ for each fixed $x_2$, and convex in $x_2$ for each fixed $x_1$. Is it convex on $\R^2$?

Compute the Hessian at $(1,1)$ and check its eigenvalues (or its determinant).

$\hess f=\begin{pmatrix}2x_2^2&4x_1x_2\\4x_1x_2&2x_1^2\end{pmatrix}$. At $(1,1)$ it is $\begin{pmatrix}2&4\\4&2\end{pmatrix}$, with determinant $4-16=-12\lt0$: eigenvalues $-2$ and $6$, indefinite. So no. A concrete bad chord: take $\x=(2,0)$, $\y=(0,2)$: $f(\x)=f(\y)=0$ but the midpoint $(1,1)$ has $f=1\gt0$. Convex in each variable separately does not mean convex.

For log-sum-exp in two variables, $f(x_1,x_2)=\log(e^{x_1}+e^{x_2})$, compute the Hessian at $\x=(\log3,\ 0)$ and give its largest eigenvalue.

$\p=(e^{x_1},e^{x_2})/(e^{x_1}+e^{x_2})=(3,1)/4$. Use $\hess f=\operatorname{diag}(\p)-\p\p^\top$.

$\p=(0.75,0.25)$. $\hess f=\begin{pmatrix}0.75-0.5625&-0.1875\\-0.1875&0.25-0.0625\end{pmatrix}=\begin{pmatrix}0.1875&-0.1875\\-0.1875&0.1875\end{pmatrix}$. Eigenvalues $0$ (eigenvector $(1,1)$) and $0.375$ (eigenvector $(1,-1)$). PSD but not PD: along $(1,1)$, $f(\x+t(1,1))=f(\x)+t$ is a straight line.

$f(x)=-\log x$ on $(0,\infty)$. Find the tangent line at $x=1$ and compute the gap $f(3)-\big(\text{tangent at }3\big)$.

$f(1)=0$, $f'(x)=-1/x$ so $f'(1)=-1$. The tangent is $0-1\cdot(y-1)$.

Tangent: $y\mapsto1-y$, which at $y=3$ is $-2$. $f(3)=-\log3\approx-1.0986$. Gap $=-1.0986-(-2)=2-\log3\approx0.9014\gt0$, as the first-order characterization guarantees for this convex function ($f''=1/x^2\gt0$).

  • To check a specific $C^2$ function: Hessian, then PSD at every point
  • To prove a general statement: start from $f(\y)\ge f(\x)+\grad f(\x)^\top(\y-\x)$
  • In proofs, name your trick: "walk a fraction $\theta$", "tangent at the mixed point", "use it from both ends", "restrict to a line"
  • Checking the Hessian at a single point and calling the function convex
  • Confusing "convex in each variable separately" with "convex"
  • Forgetting the domain must be open (for derivatives) and convex (for the segment)
  1. For $f\in C^1$: convex $\iff$ $f(\y)\ge f(\x)+\grad f(\x)^\top(\y-\x)$ $\iff$ $(\grad f(\y)-\grad f(\x))^\top(\y-\x)\ge0$.
  2. For $f\in C^2$: convex $\iff$ $\hess f(\x)\succeq0$ for every $\x$; so $\tfrac12\x^\top A\x-\b^\top\x$ is convex iff $A\succeq0$.
  3. Use the Hessian test to verify, the tangent inequality to prove; log-sum-exp is convex by Cauchy–Schwarz.

In the proof "chord $\Rightarrow$ tangent", where does the gradient come from?

From the Hessian being PSD
No second derivatives are used in this direction. Look at what happens as $\theta\to0^+$.
From the limit of the difference quotient $\frac{f(\x+\theta(\y-\x))-f(\x)}{\theta}$ as $\theta\to0^+$
That limit is the directional derivative $\grad f(\x)^\top(\y-\x)$.
From the mean value theorem applied to the chord
A mean value theorem gives a gradient at an unknown middle point, but we need it at $\x$ itself.

$f\in C^2(\R^n)$ has $\hess f(\0)\succ0$ and $\grad f(\0)=\0$. What can you conclude?

$f$ is convex, so $\0$ is the global minimizer
PSD at one point only describes $f$ near that point.
$\0$ is a strict local minimizer; nothing global follows
That is the second-order sufficient condition (Part 3). For example $(x^2-1)^2$ has $f''(\pm1)=8\gt0$ but is not convex.
$f$ is strictly convex near $\0$ and so on all of $\R^n$
Local curvature doesn't spread to the whole space.

In one variable, gradient monotonicity says…

$f$ is nondecreasing
$x^2$ is convex but decreasing for $x\lt0$. Monotonicity is about $f'$, not $f$.
$f'(x)\ge0$ everywhere
That would make $f$ nondecreasing. Write out $(f'(y)-f'(x))(y-x)\ge0$.
$f'$ is nondecreasing: the slope never goes down as $x$ increases
$(f'(y)-f'(x))(y-x)\ge0$ means $f'(y)\ge f'(x)$ whenever $y\gt x$.

The quadratic $\tfrac12\x^\top A\x-\b^\top\x$ with $A=\begin{pmatrix}1&2\\2&1\end{pmatrix}$ is…

convex, since every entry of $A$ is positive
Positive entries don't imply PSD. Compute the determinant.
not convex, since $A$ has eigenvalue $-1$
$\det A=1-4=-3\lt0$, so the eigenvalues ($3$ and $-1$) have opposite signs: indefinite.
convex, since $A$ is symmetric
Symmetric matrices can be indefinite.

Local means global, and the convexity hierarchy

For a convex function, every local minimizer is a global minimizer, and (for $f\in C^1$) $\grad f(\x^\star)=\0$ is both necessary and sufficient for global optimality. Strict convexity adds uniqueness; strong convexity adds existence.

This is the theorem that lets you say gradient descent found the answer rather than an answer, and the existence–uniqueness proof is, in the book's words, a near-guaranteed exam question.

In a single bowl, "nothing nearby is lower" already means "nothing anywhere is lower": any lower point elsewhere would drag the rim down between you and it.

Try it

Click anywhere on the map to drop a ball; it rolls downhill by gradient descent. Orange tint marks where the Hessian is not PSD. On the non-convex landscapes, find starts that end at different heights. On the two convex ones, try to make two balls end at different heights: you can't. On the flat valley, balls stop at different points with the same value.

Let $f:\R^n\to\R$ be convex.

  1. Every local minimizer of $f$ is a global minimizer.
  2. If moreover $f\in C^1$, then $\grad f(\x^\star)=\0$ $\iff$ $\x^\star$ is a global minimizer.
  3. The set of global minimizers is convex (possibly empty).

Prove parts 1 and 2 (Tutorial 1 Problem 13).

  1. Let $\x^\star$ be a local minimizer: $f(\x^\star)\le f(\z)$ for all $\z$ with $\norm{\z-\x^\star}\lt\varepsilon$. Suppose, for contradiction, some $\y$ has $f(\y)\lt f(\x^\star)$.

    Proof by contradiction: assume a strictly lower point exists somewhere.

  2. Walk from $\x^\star$ towards $\y$: $\x_\theta=(1-\theta)\x^\star+\theta\y$, so $\norm{\x_\theta-\x^\star}=\theta\norm{\y-\x^\star}$. Choose $\theta=\min\{\tfrac12,\ \varepsilon/(2\norm{\y-\x^\star})\}\in(0,1)$, so $\x_\theta$ is inside the $\varepsilon$-ball.

    We need a point that is both on the segment and close to $\x^\star$; a small enough step achieves both.

  3. Convexity: $f(\x_\theta)\le(1-\theta)f(\x^\star)+\theta f(\y)\lt(1-\theta)f(\x^\star)+\theta f(\x^\star)=f(\x^\star)$.

    The strict $\lt$ uses $f(\y)\lt f(\x^\star)$ and $\theta\gt0$.

  4. So a point in the $\varepsilon$-ball is strictly lower than $\x^\star$: contradiction. Hence no such $\y$ exists, and $\x^\star$ is global.

    The chord from $\x^\star$ down to $\y$ drags $f$ below $f(\x^\star)$ arbitrarily close to $\x^\star$.

  5. Part 2. ($\Leftarrow$) If $\grad f(\x^\star)=\0$, the tangent inequality gives $f(\y)\ge f(\x^\star)+\0^\top(\y-\x^\star)=f(\x^\star)$ for every $\y$: global. ($\Rightarrow$) A global minimizer is a local one, so the first-order necessary condition (Part 3) gives $\grad f(\x^\star)=\0$.

    This is the upgrade convexity buys: the necessary condition of Part 3 becomes sufficient, globally.

Part 3 is one line: the minimizers are the sublevel set $\{\x:f(\x)\le\min f\}$, which is convex by Chapter 4.1. So a convex function can have no minimizer ($e^x$), exactly one ($x^2$), or a whole convex set of them ($(x_1+x_2-1)^2$ has a whole line), but never, say, exactly two.

Why this matters for algorithms. Gradient descent stops when $\norm{\grad f(\x_k)}\approx0$. For general $f$ that only certifies a stationary point, which might be a saddle or a poor local minimum. For convex $f$ it certifies (approximately) the global minimum. Every convergence theorem from Lecture 8 onwards assumes convexity for exactly this reason.

The hierarchy: strictly and strongly convex

Convexity alone allows flat stretches (the chord can lie on the graph) and allows no minimizer at all. Two stronger notions fix these.

Let $C$ be convex and $f:C\to\R$.

  • $f$ is strictly convex if $f(\theta\x+(1-\theta)\y)\lt\theta f(\x)+(1-\theta)f(\y)$ for all $\x\ne\y$ and $\theta\in(0,1)$: chords lie strictly above the graph.
  • $f\in C^1$ is $\mu$-strongly convex ($\mu\gt0$) if for all $\x,\y\in C$ $$f(\y)\ \ge\ f(\x)+\grad f(\x)^\top(\y-\x)+\frac\mu2\norm{\y-\x}^2.$$ For $f\in C^2$ this is equivalent to $\hess f(\x)\succeq\mu I$ for all $\x$, i.e. $\lambda_{\min}(\hess f(\x))\ge\mu$ everywhere.
  • $f$ is quasiconvex if every sublevel set $\{\x:f(\x)\le\alpha\}$ is convex; equivalently $f(\theta\x+(1-\theta)\y)\le\max\{f(\x),f(\y)\}$.

Strict hierarchy: strongly convex $\subsetneq$ strictly convex $\subsetneq$ convex $\subsetneq$ quasiconvex.

Picture it. Strong convexity says that at every point there is not just a tangent plane below $f$, but a whole upward parabola of curvature $\mu$ below it. The function curves up at least as fast as $\tfrac\mu2\norm{\x}^2$, everywhere. A useful equivalent: $f$ is $\mu$-strongly convex iff $f(\x)-\tfrac\mu2\norm{\x}^2$ is still convex (for $C^2$ functions: $\hess f-\mu I\succeq0$).

Examples. $\tfrac12\x^\top A\x-\b^\top\x$ with $A\succ0$ is $\lambda_{\min}(A)$-strongly convex; ridge regression is $\lambda$-strongly convex. The worked example of Chapter 4.1 showed $x^2$ is $2$-strongly convex.

Separating exampleWhy
Strictly but not strongly convex: $x^4$$f''(x)=12x^2\gt0$ for $x\ne0$, but $f''(0)=0$, so no $\mu\gt0$ works near 0.
Strictly but not strongly convex: $e^x$$f''=e^x\gt0$, but $\inf_xe^x=0$. It also has no minimizer.
Convex but not strictly convex: $|x|$, affine functionsThe chord lies on the graph wherever $f$ is linear.
Quasiconvex but not convex: $\sqrt{|x|}$Sublevel sets $[-\alpha^2,\alpha^2]$ are intervals, but $f''\lt0$ for $x\gt0$.
  1. (a) If $f$ is strictly convex, it has at most one global minimizer.
  2. (b) If $f:\R^n\to\R$ is $\mu$-strongly convex, a global minimizer exists and is unique. Moreover, $f(\x)\ge f(\x^\star)+\tfrac\mu2\norm{\x-\x^\star}^2$ for all $\x$ ([Y] Thm. 2.1.8).

Prove the theorem (Tutorial 1 Problem 17). The book's Examtip: be able to reproduce this chain verbatim.

  1. (a) Suppose $\x^\star\ne\y^\star$ are both global minimizers with value $f^\star$. At the midpoint $\z=\tfrac12\x^\star+\tfrac12\y^\star$, strict convexity gives $f(\z)\lt\tfrac12f^\star+\tfrac12f^\star=f^\star$. That contradicts $f^\star$ being the minimum value.

    Two minimizers would make the chord between them flat at height $f^\star$; strictly convex means the graph is strictly below that chord.

  2. (b) Coercivity. Fix any $\x_0$, let $\g_0=\grad f(\x_0)$ and $R=\norm{\x-\x_0}$. Strong convexity and Cauchy–Schwarz ($\g_0^\top(\x-\x_0)\ge-\norm{\g_0}R$) give $$f(\x)\ \ge\ f(\x_0)-\norm{\g_0}R+\tfrac\mu2R^2.$$

    The worst the linear term can do is pull down by $\norm{\g_0}R$; the quadratic term pushes up by $\frac\mu2R^2$.

  3. As $R\to\infty$ the $R^2$ term wins (because $\mu\gt0$), so $f(\x)\to\infty$ as $\norm{\x}\to\infty$: $f$ is coercive. Concretely, $f(\x)\gt f(\x_0)$ whenever $R\gt2\norm{\g_0}/\mu$.

    $-\norm{\g_0}R+\frac\mu2R^2=R(\frac\mu2R-\norm{\g_0})\gt0$ exactly when $R\gt2\norm{\g_0}/\mu$.

  4. Compact sublevel set. $S=\{\x:f(\x)\le f(\x_0)\}$ is nonempty ($\x_0\in S$), closed ($f$ is continuous) and bounded (inside the ball of radius $2\norm{\g_0}/\mu$ around $\x_0$). Closed and bounded in $\R^n$ means compact (Heine–Borel).

    Weierstrass needs compactness; $\R^n$ itself isn't compact, so we cut down to a sublevel set.

  5. Weierstrass. The continuous $f$ attains its minimum on the compact $S$ at some $\x^\star$. It is global: points outside $S$ have $f\gt f(\x_0)\ge f(\x^\star)$.

    Existence. Note that convexity was used only to get coercivity.

  6. Uniqueness. For $\x\ne\y$, the extra term $\frac\mu2\norm{\y-\x}^2\gt0$ makes the tangent inequality strict, which gives strict convexity (run the "tangent at the mixed point" proof of Chapter 4.2 with strict inequalities). Part (a) applies.

    Strong $\Rightarrow$ strict $\Rightarrow$ at most one minimizer.

  7. Quadratic growth. Put $\x=\x^\star$ in the strong convexity inequality and use $\grad f(\x^\star)=\0$: $f(\y)\ge f(\x^\star)+\frac\mu2\norm{\y-\x^\star}^2$.

    The function rises at least quadratically away from its minimizer, which is what later gives fast convergence.

Careful: each hypothesis is doing work. $e^x$ is strictly convex with no minimizer (strict convexity gives uniqueness, not existence). $(x_1+x_2-1)^2$ is convex with infinitely many minimizers. And coercivity alone, without convexity, gives existence but not uniqueness: Tutorial 1 Problem 19's $x_1^4-4x_1x_2+x_2^4$ has two global minimizers.
Go deeper: pseudoconvexity, and a slicker proof for the quartic

Pseudoconvex ($f\in C^1$): $\grad f(\x)^\top(\y-\x)\ge0\Rightarrow f(\y)\ge f(\x)$. Put $\grad f(\x)=\0$: every stationary point is a global minimizer, convexity's most useful property, without being convex. Example: $x+x^3$ has $f'\gt0$, so $f'(x)(y-x)\ge0$ forces $y\ge x$ and then $f(y)\ge f(x)$: pseudoconvex, yet $f''=6x\lt0$ for $x\lt0$. Quasiconvex does not imply pseudoconvex: $x^3$ is increasing, so its sublevel sets are intervals, but $x=0$ is stationary and $f(-1)\lt f(0)$.

The quartic $h=x_1^4-4x_1x_2+x_2^4$ in one line. Since $(x^2-1)^2\ge0$, $x^4\ge2x^2-1$. So $x_1^4+x_2^4\ge2(x_1^2+x_2^2)-2\ge4x_1x_2-2$ (using $x_1^2+x_2^2\ge2x_1x_2$). Hence $h\ge-2$ everywhere, with equality iff $x_1^2=x_2^2=1$ and $x_1=x_2$: the minimizers are exactly $(1,1)$ and $(-1,-1)$.

How many global minimizers does $f(x_1,x_2)=(x_1+x_2-2)^2$ have?

$f\ge0$. Where is $f=0$?

$f=0$ on the whole line $x_1+x_2=2$, and $f\ge0$ everywhere, so every point of that line is a global minimizer: infinitely many. $f$ is convex (Hessian $\begin{pmatrix}2&2\\2&2\end{pmatrix}$, eigenvalues 0 and 4) but not strictly convex, so uniqueness fails; the minimizer set is a line, a convex set.

What is the largest $\mu$ for which $f(x)=x^2+e^x$ is $\mu$-strongly convex on $\R$?

For $C^2$ functions of one variable: $\mu$-strongly convex iff $f''(x)\ge\mu$ for all $x$. Find $\inf_xf''(x)$.

$f''(x)=2+e^x\gt2$, and $\inf_x(2+e^x)=2$ (approached as $x\to-\infty$, never attained). So $f''\ge\mu$ everywhere iff $\mu\le2$: the largest is $\mu=2$. The infimum need not be attained for this to work.

$f$ is $2$-strongly convex with minimizer $\x^\star$ and $f(\x^\star)=1$. What is the best lower bound on $f(\x)$ that the theory gives at a point with $\norm{\x-\x^\star}=3$?

Quadratic growth: $f(\x)\ge f(\x^\star)+\frac\mu2\norm{\x-\x^\star}^2$.

$f(\x)\ge1+\frac22\cdot9=10$. It is attained by $f(\x)=1+\norm{\x-\x^\star}^2$, so it can't be improved.

In the existence proof, $f$ is $0.5$-strongly convex and at the base point $\norm{\grad f(\x_0)}=3$. Give the radius $R$ (around $\x_0$) of the ball that the proof shows must contain every global minimizer.

$f(\x)\ge f(\x_0)-\norm{\g_0}R+\frac\mu2R^2$, and this exceeds $f(\x_0)$ once $R\gt2\norm{\g_0}/\mu$.

$R=2\cdot3/0.5=12$. Outside that ball $f(\x)\gt f(\x_0)$, so no minimizer can be there: the minimizer is within distance 12 of $\x_0$.

(Tutorial 1 Problem 19.) $h(x,y)=x^4-4xy+y^4$ is not convex. Show it has a global minimum and find its value.

Coercivity: $-4xy\ge-2(x^2+y^2)$ and $x^4+y^4\ge\frac12(x^2+y^2)^2$, so $h\ge\frac{r^4}2-2r^2\to\infty$. Then a global minimizer exists and is a critical point: solve $\grad h=\0$.

With $r^2=x^2+y^2$: $h\ge\frac{r^4}{2}-2r^2\to\infty$, so $h$ is coercive and (Weierstrass on a sublevel set) has a global minimizer, which must satisfy $\grad h=(4x^3-4y,\,4y^3-4x)=\0$: $y=x^3$, $x=y^3$, so $x=x^9$, $x\in\{0,\pm1\}$. Critical points $(0,0)$, $(1,1)$, $(-1,-1)$ with values $0,-2,-2$. The minimum value is $-2$, attained at two points. ($\hess h(0,0)$ has eigenvalues $\pm4$: a saddle. So $h$ is not convex.)

  • For convex $C^1$ $f$, certify a global minimizer by checking $\grad f(\x^\star)=\0$
  • Prove existence by: strong convexity → coercive → compact sublevel set → Weierstrass
  • Find $\mu$ as $\inf_\x\lambda_{\min}(\hess f(\x))$ over the whole domain
  • Saying "strictly convex, so a minimizer exists" ($e^x$)
  • Using $f''\gt0$ everywhere as proof of strong convexity ($x^4$, $e^x$)
  • Forgetting that a non-convex function's stationary point might be a saddle or a poor local minimum
  1. Convex: local minimizer = global minimizer, and for $f\in C^1$, $\grad f(\x^\star)=\0\iff$ global minimizer; minimizers form a convex set.
  2. Strict convexity gives at most one minimizer; $\mu$-strong convexity ($\hess f\succeq\mu I$) gives exactly one, plus $f(\x)-f^\star\ge\frac\mu2\norm{\x-\x^\star}^2$.
  3. The existence chain: strongly convex → coercive → compact sublevel set → Weierstrass → exists; strict → unique.

Gradient descent on a convex $C^1$ function stops at $\x_k$ with $\grad f(\x_k)=\0$. Then $\x_k$ is…

a stationary point, possibly a saddle
For general $f$ yes, but here $f$ is convex. What does the tangent inequality say at $\x_k$?
a global minimizer, though maybe not the only one
Tangent inequality with zero gradient: $f(\y)\ge f(\x_k)$ for all $\y$. Uniqueness needs strict convexity.
the unique global minimizer
Only if $f$ is strictly convex. $(x_1+x_2-1)^2$ has a whole line of minimizers.

Which function is strictly convex but has no minimizer?

$x^4$
$x^4$ has its minimizer at 0.
$|x|$
$|x|$ is minimized at 0, and it isn't strictly convex.
$e^x$
$e^x\gt0$ with $\inf e^x=0$ never attained. Strictly convex, not strongly: $f''=e^x\to0$.

In the existence proof for strongly convex $f$, why can't we apply Weierstrass on $\R^n$ directly?

$\R^n$ is not compact (it is unbounded), so we restrict to a sublevel set that is closed and bounded
Coercivity is exactly what makes the sublevel set bounded.
$f$ might not be continuous on $\R^n$
$f\in C^1$, so it is continuous.
Weierstrass only works in one dimension
It works on any compact subset of $\R^n$.

$f\in C^2(\R^n)$ with $\hess f(\x)\succ0$ for every $\x$. Which is guaranteed?

$f$ is strongly convex
Not without a uniform lower bound $\mu$ on the eigenvalues: think of $e^x$.
$f$ is strictly convex, so it has at most one minimizer
Positive definite everywhere gives strict convexity (strict tangent inequality). Existence is not guaranteed.
$f$ has exactly one minimizer
Existence can fail: $e^{x_1}+x_2^2$ has PD Hessian everywhere and no minimizer.

Smoothness, the quadratic sandwich, and the condition number

An $L$-smooth function curves up at most as fast as $\frac L2\norm{\x}^2$; a $\mu$-strongly convex one curves up at least as fast as $\frac\mu2\norm{\x}^2$. A function with both is sandwiched between two parabolas at every point.

The convergence theorems of Lectures 8–10 are stated for exactly these classes, written $\mathcal F_L^{1,1}$ and $\mathcal S_{\mu,L}^{1,1}$, and their speed depends on one number, $\kappa=L/\mu$.

A road with a speed bump limit and a minimum incline: the curvature can't exceed $L$ (no sudden sharp turns), and can't drop below $\mu$ (no flat stretches where you'd stall).

$f\in C^1(\R^n)$ is $L$-smooth if its gradient is $L$-Lipschitz: $$\norm{\grad f(\x)-\grad f(\y)}\le L\norm{\x-\y}\qquad\text{for all }\x,\y.$$ For $f\in C^2$ this is equivalent to $\norm{\hess f(\x)}_2\le L$ for all $\x$, i.e. every eigenvalue of every Hessian lies in $[-L,L]$. For convex $f\in C^2$ it becomes $0\preceq\hess f(\x)\preceq LI$.

If $f$ is $L$-smooth then, for all $\x,\y$ (no convexity needed), $$f(\y)\ \le\ f(\x)+\grad f(\x)^\top(\y-\x)+\frac L2\norm{\y-\x}^2.$$

Compare with strong convexity, the same expression with $\mu$ and $\ge$. Putting them together:

If $f$ is $\mu$-strongly convex and $L$-smooth, then for all $\x,\y$: $$\frac\mu2\norm{\y-\x}^2\ \le\ f(\y)-f(\x)-\grad f(\x)^\top(\y-\x)\ \le\ \frac L2\norm{\y-\x}^2.$$ At the minimizer ($\grad f(\x^\star)=\0$): $\ \frac\mu2\norm{\y-\x^\star}^2\le f(\y)-f^\star\le\frac L2\norm{\y-\x^\star}^2$. Necessarily $\mu\le L$.

The middle quantity is "how far $f$ rises above its tangent". In one variable, Taylor says it is about $\frac12f''\cdot(y-x)^2$, so the sandwich is the statement $\mu\le f''\le L$ everywhere. For a quadratic $\tfrac12\x^\top A\x-\b^\top\x$, the middle quantity is exactly $\frac12(\y-\x)^\top A(\y-\x)$, and the Rayleigh bounds of Part 0b give the sandwich with $\mu=\lambda_{\min}(A)$, $L=\lambda_{\max}(A)$.

Try it

Drag along the plot to move the anchor $\x$. The green parabola (curvature $\mu$) must stay below $f$ and the purple one (curvature $L$) above it. Raise $\mu$ until the lower bound breaks somewhere, then lower $L$ until the upper bound breaks. "Use the tightest" sets $\mu=\inf f''$, $L=\sup f''$. Compare: the logistic loss (no $\mu\gt0$ works), $x^4$ (no $L$ works on all of $\R$), and the non-convex function (an $L$ exists, but no lower parabola with $\mu\ge0$).

  • $\mathcal F_L^{1,1}(\R^n)$: convex, continuously differentiable functions with $L$-Lipschitz gradient.
  • $\mathcal S_{\mu,L}^{1,1}(\R^n)$: the functions in $\mathcal F_L^{1,1}$ that are also $\mu$-strongly convex.
  • For $f\in\mathcal S_{\mu,L}^{1,1}$, the condition number is $\kappa=Q_f=L/\mu\ge1$.

For $f\in C^2$: $f\in\mathcal F_L^{1,1}\iff0\preceq\hess f\preceq LI$, and $f\in\mathcal S_{\mu,L}^{1,1}\iff\mu I\preceq\hess f\preceq LI$ everywhere.

Reading the symbols
$\mathcal F$
convex functions; $\mathcal S$: strongly convex
superscript $1,1$
once continuously differentiable ($1$), with the $1$st derivative Lipschitz ($1$)
subscripts $\mu,L$
the curvature floor and ceiling

Place the fit-the-line loss $L(w,c)$ and $f(x,y)=3x^2+2xy+3y^2$ in the hierarchy, and compute their $\mu$, $L$ and $\kappa$.

  1. Fit-the-line: $L(\z)=\tfrac12\z^\top A\z-\b^\top\z+38$ with $A=\begin{pmatrix}28&12\\12&6\end{pmatrix}$, constant Hessian $A$.

    For a quadratic, $\mu$ and $L$ are just the extreme eigenvalues of $A$.

  2. Eigenvalues: trace $34$, determinant $168-144=24$, so $\lambda=17\pm\sqrt{289-24}=17\pm16.279$, i.e. $\mu\approx0.721$, $L\approx33.28$, and $\kappa\approx46.1$.

    For a $2\times2$ symmetric matrix, $\lambda=\frac{\mathrm{tr}}2\pm\sqrt{(\frac{\mathrm{tr}}2)^2-\det}$.

  3. So $L(w,c)\in\mathcal S_{0.721,\,33.28}^{1,1}$: it has a unique minimizer $(w,c)=(1.5,\,1/3)$ and is a long, thin valley, 46 times steeper across than along.

    Level sets have axis ratio $\sqrt\kappa\approx6.8$ (Part 0b).

  4. $f(x,y)=3x^2+2xy+3y^2=\tfrac12\x^\top\begin{pmatrix}6&2\\2&6\end{pmatrix}\x$. Eigenvalues $6\pm2$: $\mu=4$, $L=8$, $\kappa=2$.

    Careful with the factor $\tfrac12$: the matrix entries are twice the coefficients of $x^2$ and $y^2$, and the off-diagonal is the $xy$ coefficient.

The bridge to Part 5

Part 5 runs gradient descent on convex quadratics and shows the error shrinks by a factor about $\frac{\kappa-1}{\kappa+1}$ per step (with the best constant step). For $\kappa=2$ that's $\tfrac13$ per step: fast. For the fit-the-line loss, $\kappa\approx46.1$ gives about $0.958$ per step: slow zig-zagging along the valley. Lectures 9–10 prove the same kind of rates for all of $\mathcal F_L^{1,1}$ (an $O(1/k)$ rate) and $\mathcal S_{\mu,L}^{1,1}$ (a linear rate governed by $\kappa$), using the sandwich as their main tool. In every one of those proofs, $L$ sets the safe step size ($h\le1/L$ or $2/(\mu+L)$) and $\mu$ sets how much progress each step guarantees.

Go deeper: proof of the upper bound, and the gradient sandwich

Proof of the quadratic upper bound. With the integral form of Taylor, $f(\y)-f(\x)-\grad f(\x)^\top(\y-\x)=\int_0^1\big(\grad f(\x+t(\y-\x))-\grad f(\x)\big)^\top(\y-\x)\,dt$. By Cauchy–Schwarz and the Lipschitz condition the integrand is at most $Lt\norm{\y-\x}^2$, and $\int_0^1Lt\,dt=\frac L2$. Part 6 uses this as the Descent Lemma: with $\y=\x-\frac1L\grad f(\x)$ it gives $f(\y)\le f(\x)-\frac1{2L}\norm{\grad f(\x)}^2$.

Why $\norm{\hess f}\le L$ and not just $\hess f\preceq LI$. Lipschitz bounds the gradient's change in both directions, so very negative curvature is ruled out too: $-\tfrac{M}2x^2$ with $M\gt L$ is not $L$-smooth, even though its Hessian is $\preceq LI$. For convex $f$ the lower side is automatic ($\hess f\succeq0$), which is why $0\preceq\hess f\preceq LI$ characterizes $\mathcal F_L^{1,1}$ ([Y] Thm. 2.1.6).

The gradient sandwich. For $f\in\mathcal S_{\mu,L}^{1,1}$ ([Y] (2.1.7) and (2.1.19) with one point at $\x^\star$): $$\frac1{2L}\norm{\grad f(\x)}^2\ \le\ f(\x)-f^\star\ \le\ \frac1{2\mu}\norm{\grad f(\x)}^2.$$ So a small gradient really does certify a small optimality gap, but only when $\mu\gt0$: this is the quantitative version of "local = global". [Y] Thm. 2.1.5 lists further equivalent forms of $f\in\mathcal F_L^{1,1}$, including co-coercivity $\big(\grad f(\x)-\grad f(\y)\big)^\top(\x-\y)\ge\frac1L\norm{\grad f(\x)-\grad f(\y)}^2$, which Lecture 9 uses.

Find the best (largest) $\mu$ and best (smallest) $L$ for $f(x)=x^2+\sin x$ on $\R$.

$f''(x)=2-\sin x$. What range does $\sin x$ cover?

$f''=2-\sin x\in[1,3]$, and both ends are attained ($x=\pi/2$ and $x=-\pi/2$). So $\mu=1$, $L=3$, $\kappa=3$: $f\in\mathcal S^{1,1}_{1,3}$.

$f(\x)=\tfrac12\x^\top A\x$ with $A=\begin{pmatrix}5&2\\2&2\end{pmatrix}$. Compute $\kappa=L/\mu$.

Trace 7, determinant $10-4=6$. The eigenvalues are the roots of $\lambda^2-7\lambda+6$.

$\lambda^2-7\lambda+6=(\lambda-1)(\lambda-6)$, so $\mu=1$, $L=6$, $\kappa=6$.

$f$ is $4$-smooth. At $\x$, $f(\x)=2$ and $\grad f(\x)=(1,-1)$. Using the quadratic upper bound, how large can $f(\y)$ be at $\y=\x-\frac14\grad f(\x)$?

$\y-\x=-\frac14\g$. Substitute: $f(\y)\le f(\x)-\frac14\norm{\g}^2+\frac42\cdot\frac1{16}\norm{\g}^2$.

$\norm{\g}^2=2$. $f(\y)\le2-\tfrac14\cdot2+2\cdot\tfrac1{16}\cdot2=2-0.5+0.25=1.75$. In general, a step of $1/L$ decreases $f$ by at least $\frac1{2L}\norm{\g}^2=\frac18\cdot2=0.25$.

Classify the logistic loss $f(x)=\log(1+e^{-x})$.

$f'(x)=-\frac1{1+e^x}$ and $f''(x)=\sigma(x)(1-\sigma(x))$ with $\sigma(x)=\frac1{1+e^{-x}}\in(0,1)$. What are $\sup f''$ and $\inf f''$?

$f''=\sigma(1-\sigma)\in(0,\tfrac14]$: positive (convex, even strictly), at most $\tfrac14$ (attained at $x=0$, where $\sigma=\tfrac12$), and tending to 0 as $|x|\to\infty$. So $f\in\mathcal F^{1,1}_{1/4}$ but $\inf f''=0$: no $\mu\gt0$, not strongly convex. (It also has no minimizer: $f\to0$ as $x\to\infty$.)

  • Read off $\mu$ and $L$ as the inf and sup over $\x$ of $\lambda_{\min}$ and $\lambda_{\max}$ of $\hess f(\x)$
  • Remember the $\tfrac12$ when converting $ax^2+bxy+cy^2$ to $\tfrac12\x^\top A\x$
  • Expect slow gradient descent when $\kappa$ is large
  • Assuming polynomials are $L$-smooth on $\R^n$ ($x^4$ isn't)
  • Thinking $L$-smooth requires convexity (the upper bound holds without it)
  • Taking $\mu=\min f''$ at the minimizer only; it must hold everywhere
  1. $L$-smooth: $\norm{\grad f(\x)-\grad f(\y)}\le L\norm{\x-\y}$, giving $f(\y)\le f(\x)+\grad f(\x)^\top(\y-\x)+\frac L2\norm{\y-\x}^2$.
  2. $\mathcal F_L^{1,1}$ = convex and $L$-smooth ($0\preceq\hess f\preceq LI$); $\mathcal S_{\mu,L}^{1,1}$ adds $\mu$-strong convexity ($\mu I\preceq\hess f\preceq LI$): $f$ is sandwiched between two parabolas.
  3. $\kappa=L/\mu$; for quadratics $\kappa=\lambda_{\max}(A)/\lambda_{\min}(A)$, and it governs gradient descent's speed (Part 5).

$f(x)=x^4$ on $\R$ belongs to…

$\mathcal S_{\mu,L}^{1,1}$ for suitable $\mu,L$
Is there a $\mu\gt0$ with $12x^2\ge\mu$ at $x=0$?
$\mathcal F_L^{1,1}$ for some $L$
Is $12x^2$ bounded on $\R$?
neither: it is convex but not $L$-smooth on $\R$ (and not strongly convex)
$f''=12x^2$ is unbounded, so no $L$; $f''(0)=0$, so no $\mu$.

For $f\in\mathcal S_{\mu,L}^{1,1}$ with minimizer $\x^\star$, which bound on $f(\x)-f^\star$ is right?

$\frac\mu2\norm{\x-\x^\star}^2\le f(\x)-f^\star\le\frac L2\norm{\x-\x^\star}^2$
The sandwich at $\x^\star$, where the gradient term vanishes.
$\frac L2\norm{\x-\x^\star}^2\le f(\x)-f^\star\le\frac\mu2\norm{\x-\x^\star}^2$
$\mu\le L$, so this has the bounds the wrong way round.
$f(\x)-f^\star\le\mu\norm{\x-\x^\star}$
The bounds are quadratic in the distance, not linear.

$f(\x)=\tfrac12\x^\top A\x-\b^\top\x$ with $A\succeq0$ singular. Then…

$f\in\mathcal S_{\mu,L}^{1,1}$ with $\mu=\lambda_{\min}(A)$
$\lambda_{\min}(A)=0$ here, and strong convexity needs $\mu\gt0$.
$f\in\mathcal F_L^{1,1}$ with $L=\lambda_{\max}(A)$, but it is not strongly convex
Convex and $L$-smooth; along the null space of $A$ it is flat or linear.
$f$ is not convex
$\hess f=A\succeq0$, so it is convex.

Which statement about $L$-smoothness is true?

It requires $f$ to be convex
The definition only constrains how fast $\grad f$ changes. The widget's non-convex example is 3-smooth.
It gives a quadratic upper bound on $f$ around every point, with or without convexity
That is the Descent Lemma, the key to safe step sizes.
It means $\hess f\succeq LI$
That would be a lower bound on curvature, like strong convexity. Smoothness bounds curvature from above.

$f$ is convex and $C^1$ on $\R^n$, and $\grad f(\x^\star)=\0$. Which statement is the strongest correct conclusion?

$\x^\star$ is a local minimizer
True, but you can say much more for convex $f$.
$\x^\star$ is a global minimizer
The tangent inequality at $\x^\star$ with zero gradient gives $f(\y)\ge f(\x^\star)$ for all $\y$.
$\x^\star$ is the unique global minimizer
Uniqueness needs strict convexity.

The Hessian of a $C^2$ function at every point has eigenvalues in $[2,10]$. Then $f$ is…

convex but not necessarily strongly convex
A uniform positive lower bound on the eigenvalues is exactly strong convexity.
in $\mathcal S^{1,1}_{2,10}$, with $\kappa\le5$
$2I\preceq\hess f\preceq10I$ everywhere.
only locally convex near its minimizer
The bound holds at every point, so the conclusion is global.

Why can $f(x)=x^4-2x^2+0.5x$ have a local minimizer that isn't global, while $x^4+0.5x$ can't?

Because $x^4+0.5x$ has no critical points
$4x^3+0.5=0$ has the solution $x=-0.5$.
Because $x^4+0.5x$ is convex ($f''=12x^2\ge0$) while the first has $f''(0)=-4\lt0$
Convexity forbids "lower point elsewhere"; the $-2x^2$ term creates a second valley.
Because quartics always have a unique minimizer
$x^4-2x^2$ is a quartic with two minimizers.

To show $f(\x)=\norm{A\x-\b}^2+\lambda\norm{\x}^2$ ($\lambda\gt0$) is strongly convex, the quickest route is…

the chord definition with $\theta=\tfrac12$
One value of $\theta$ doesn't prove anything for all pairs, and the algebra is long.
$\hess f=2A^\top A+2\lambda I\succeq2\lambda I$
Hessian test: $\v^\top\hess f\v=2\norm{A\v}^2+2\lambda\norm{\v}^2\ge2\lambda\norm{\v}^2$, so $\mu=2\lambda$.
showing $\grad f(\x^\star)=\0$ at the minimizer
That certifies a minimizer of a convex function; it says nothing about strong convexity.

Gradient descent with a fixed step cycles on $f(x)=|x|$ (Part 1). Which class does $|x|$ fail to belong to?

Convex functions
$|x|$ is a norm, hence convex.
$\mathcal F_L^{1,1}$: it is not differentiable at 0, let alone with a Lipschitz gradient
The rates of Lectures 9–10 need $L$-smoothness; a kink breaks it.
Quasiconvex functions
Its sublevel sets $[-c,c]$ are intervals, so it is quasiconvex.

The set of global minimizers of a convex function can be…

exactly two isolated points
The minimizer set is a sublevel set of a convex function, so it is convex: the segment between two minimizers would also consist of minimizers.
empty, a single point, or an infinite convex set
$e^x$, $x^2$ and $(x_1+x_2)^2$ respectively.
only a single point
$(x_1+x_2)^2$ is minimized on a whole line.

Lecture 6 of the course: the gradient descent algorithm, why each step is the exact minimizer of a simple quadratic model, and a complete analysis on convex quadratic "bowls": which fixed step sizes work, which one is best, exact line search and its zig-zag, and the famous $\left(\frac{\kappa-1}{\kappa+1}\right)^2$ rate.

You need: Part 0b (gradient, eigenvalues, the spectral theorem, condition number $\kappa$), Part 1 (Chapter 1.3: gradient descent in one variable), Parts 3–4 (stationary points, convexity) help with the "why".

Gradient descent, and why each step is a tiny optimization problem

Gradient descent repeats one move: from where you stand, step a distance proportional to the slope, straight downhill: $\x_{k+1}=\x_k-\alpha_k\grad f(\x_k)$.

It is the workhorse of modern machine learning and the reference method every other algorithm in this course is compared against. Its step size $\alpha_k$ decides whether it crawls, converges quickly, or blows up.

Walking down a foggy hill: you can only feel the slope under your feet. You pick a stride length, step against the slope, feel again, and repeat. Too short a stride wastes time; too long and you overshoot the valley floor and land on the far slope.

Recall from Part 0b that at a point $\x$ with $\grad f(\x)\ne\0$, the direction of fastest decrease is $-\grad f(\x)$: among unit directions $\d$, the slope $\grad f(\x)^\top\d$ is smallest (equal to $-\norm{\grad f(\x)}$) exactly when $\d=-\grad f(\x)/\norm{\grad f(\x)}$ (Cauchy–Schwarz). Gradient descent simply keeps moving in that direction.

Notation for this part
$\x_k$
the $k$-th iterate (point), starting from a chosen $\x_0$
$\g_k=\grad f(\x_k)$
the gradient at the current iterate
$\alpha_k>0$
the step size (in machine learning: the "learning rate"); written $\alpha$ when it is the same every step
$\d_k$
a search direction; for gradient descent $\d_k=-\g_k$

Input: a differentiable $f$, a start $\x_0$, a tolerance $\eps>0$.

  1. For $k=0,1,2,\dots$: compute $\g_k=\grad f(\x_k)$.
  2. If $\norm{\g_k}\le\eps$, stop and return $\x_k$ (it is nearly stationary).
  3. Choose a step size $\alpha_k>0$ and set $\x_{k+1}=\x_k-\alpha_k\g_k$.

Each iteration needs one gradient and a few vector operations: $O(n)$ memory and arithmetic beyond the gradient itself. Compare (Tutorial 1 §9.2):

MethodWork per iterationMemorySpeed on a strongly convex problem
Gradient descentone gradient, $O(n)$ extra$O(n)$linear: $O(\kappa\log(1/\eps))$ iterations (this part)
Newton's methodsolve an $n\times n$ system, $O(n^3)$$O(n^2)$quadratic, near the answer (Part 10)
Quasi-Newton (BFGS)$O(n^2)$$O(n^2)$superlinear (Part 10)

With $n=10^7$ parameters (an ordinary neural network), an $n\times n$ matrix has $10^{14}$ entries: it cannot even be stored. That is why gradient descent and its relatives dominate large-scale computing, and why it matters so much to understand how fast it converges.

Each step minimizes a simple model of $f$

Why should a step of size $\alpha$ along $-\g_k$ be sensible, and not just "for very small $\alpha$"? Here is a clean answer from the lecture. Near $\x_k$, replace $f$ by the surrogate model $$m_k(\y)=f(\x_k)+\g_k^\top(\y-\x_k)+\frac1{2\alpha}\norm{\y-\x_k}^2 .$$ It is the tangent plane (first two terms) plus a round bowl of curvature $1/\alpha$ in every direction. Such a bowl is called isotropic: it curves the same amount in every direction.

The unique minimizer of $m_k$ is $\y=\x_k-\alpha\g_k$: exactly the gradient step with step size $\alpha$.

Proof. $\grad m_k(\y)=\g_k+\tfrac1\alpha(\y-\x_k)$, which is zero only at $\y=\x_k-\alpha\g_k$. The Hessian of $m_k$ is $\tfrac1\alpha I\succ0$, so $m_k$ is strictly convex and this stationary point is its unique global minimizer (Parts 3–4). $\square$

So a small $\alpha$ means a steep, narrow model bowl and a short step; a large $\alpha$ means a shallow, wide bowl and a long step. The value the model predicts at its minimizer is $$m_k(\x_k-\alpha\g_k)=f(\x_k)-\alpha\norm{\g_k}^2+\tfrac{\alpha}{2}\norm{\g_k}^2=f(\x_k)-\tfrac\alpha2\norm{\g_k}^2 .$$

When can we trust the prediction? If the model lies above $f$ everywhere (we say $m_k$ majorizes $f$), then $$f(\x_{k+1})\le m_k(\x_{k+1})=f(\x_k)-\tfrac\alpha2\norm{\g_k}^2,$$ a guaranteed decrease. This "build an upper model, minimize it, repeat" scheme is called majorize–minimize. For a quadratic $f(\x)=\tfrac12\x^\top A\x-\b^\top\x$ the second-order Taylor expansion is exact (Part 0b), so $$m_k(\y)-f(\y)=\tfrac12(\y-\x_k)^\top\Big(\tfrac1\alpha I-A\Big)(\y-\x_k),$$ which is $\ge0$ for every $\y$ exactly when $\tfrac1\alpha\ge\lambda_{\max}(A)=L$, that is, $\alpha\le 1/L$. In Part 6 the same idea, with $L$ a bound on the curvature of a general $f$, becomes the Descent Lemma.

Picture it: the surrogate is a bowl you lay over the function, touching it at $\x_k$ with the same slope. You jump to the bottom of your bowl. If your bowl is narrower than the function everywhere, the function can only be lower where you land.
Try it

The black curve is $f$, the blue parabola is the surrogate $m_k$, and the green dot is where the gradient step lands on $f$. Drag $\x_k$ around and change $\alpha$. On the bowl $x^2$ (curvature $L=2$): find the $\alpha$ that lands exactly on the minimizer in one step, the $\alpha$ where the iterate just bounces to $-x_k$, and what happens beyond it. Watch when orange shading appears (the model dips below $f$).

Choosing the step size

Three strategies run through the course:

  • Fixed step $\alpha_k=\alpha$: cheapest. Needs some knowledge of the curvature $L$. Chapter 5.2 analyses it exactly on quadratics.
  • Exact line search: choose $\alpha_k$ to minimize $f$ along the ray $\x_k-\alpha\g_k$. On a quadratic this has a closed form (Chapter 5.3); for general $f$ it is itself an optimization problem.
  • Inexact line search (Armijo, Wolfe): accept any step that decreases $f$ "enough". This is Part 6.

A fixed step can fail badly when the curvature is unbounded. On $f(x)=x^4$, gradient descent with $\alpha=0.1$ is $x_{k+1}=x_k(1-0.4x_k^2)$; from $x_0=3$ it gives $x_1=3(1-3.6)=-7.8$ and then explodes. The curvature $f''(x)=12x^2$ grows without bound, so no single $\alpha$ is safe everywhere (Tutorial 1, Problem 28). On quadratics, the curvature is the constant matrix $A$, and everything can be computed exactly. That is why this part works on quadratics.

Take one gradient step on the fit-the-line loss $L(w,c)=\tfrac12\z^\top A\z-\b^\top\z+38$, with $A=\begin{pmatrix}28&12\\12&6\end{pmatrix}$, $\b=(46,20)^\top$, from $\z_0=(0,0)^\top$ with $\alpha=0.02$. Compare the surrogate's prediction with the true new loss.

  1. $\g_0=A\z_0-\b=(-46,-20)^\top$ and $L(\z_0)=38$.

    For a quadratic, $\grad L(\z)=A\z-\b$ (Part 0b). At the origin only $-\b$ remains; the constant 38 is the loss of the line $y=0$.

  2. $\z_1=\z_0-0.02\,\g_0=(0.92,\ 0.40)^\top$: the line $y=0.92x+0.40$.

    Minus a negative gradient moves both slope and intercept up, towards the data.

  3. Surrogate prediction: $L(\z_0)-\tfrac\alpha2\norm{\g_0}^2=38-0.01\times(46^2+20^2)=38-25.16=12.84$.

    This is the formula $m_k(\x_{k+1})=f(\x_k)-\tfrac\alpha2\norm{\g_k}^2$ derived above.

  4. True loss: residuals $0.92+0.40-2=-0.68$, $1.84+0.40-3=-0.76$, $2.76+0.40-5=-1.84$, so $L(\z_1)=0.4624+0.5776+3.3856=4.4256\le12.84$.

    The largest eigenvalue of $A$ is $L\approx33.28$, so $1/L\approx0.0300$. Our $\alpha=0.02$ is below it, so the surrogate majorizes $L$ and its prediction is a guaranteed upper bound. Here the true decrease is far better than promised.

Take one gradient descent step on $f(\x)=x_1^2+3x_2^2$ from $\x_0=(1,1)^\top$ with $\alpha=0.1$. What is $\x_1$?

$\grad f=(2x_1,\ 6x_2)^\top$. Evaluate it at $\x_0$, multiply by $\alpha$, subtract.

$\g_0=(2,6)^\top$, so $\x_1=(1,1)-0.1(2,6)=(0.8,\ 0.4)^\top$. Notice the steep coordinate $x_2$ moved six times as far as it would on a round bowl: its curvature is larger.

At $\x_k$ you know $f(\x_k)=5$ and $\g_k=(3,4)^\top$. You take a gradient step with $\alpha=0.1$. What value does the surrogate $m_k$ predict at $\x_{k+1}$? (If $f$ is a quadratic with $L=4$, this is also a guaranteed upper bound on $f(\x_{k+1})$.)

$m_k(\x_{k+1})=f(\x_k)-\tfrac\alpha2\norm{\g_k}^2$, and $\norm{(3,4)}=5$.

$5-0.05\times25=3.75$. Since $1/\alpha=10\ge L=4$, the model lies above $f$, so $f(\x_{k+1})\le3.75$.

For $f(\x)=\tfrac12\x^\top A\x$ with $A=\begin{pmatrix}5&-2\\-2&2\end{pmatrix}$, what is the largest $\alpha$ for which the surrogate $m_k$ majorizes $f$ at every $\x_k$?

You need $\tfrac1\alpha I-A\succeq0$, i.e. $\tfrac1\alpha\ge\lambda_{\max}(A)$. Eigenvalues of a $2\times2$ symmetric matrix: trace $=7$, determinant $=6$.

$\lambda^2-7\lambda+6=0$ gives $\lambda\in\{1,6\}$, so $L=6$ and the largest such $\alpha$ is $1/6\approx0.1667$.

Gradient descent with fixed $\alpha=0.1$ on $f(x)=x^4$, started at $x_0=3$. What happens?

The update is $x_{k+1}=x_k(1-4\alpha x_k^2)$. It shrinks $|x|$ only when $|1-0.4x_k^2|\lt 1$. Is that true at $x_0=3$?

$|1-0.4x^2|\lt1$ needs $x^2\lt5$, i.e. $|x|\lt\sqrt5\approx2.236$. At $x_0=3$: $x_1=3(1-3.6)=-7.8$, further out, so the next factor is even larger: $x_2\approx182$, and it diverges. A smooth convex function with unbounded curvature has no globally safe fixed step.

  • Stop on a small gradient norm $\norm{\g_k}\le\eps$, not after a fixed number of steps
  • Read $\alpha$ as "one over the curvature of my model bowl"
  • Check $\alpha\le1/L$ when you want a guaranteed decrease from the surrogate argument
  • Forgetting the minus sign: $\x_k+\alpha\g_k$ climbs
  • Assuming a step that works at one point works everywhere (the $x^4$ trap)
  • Confusing the surrogate's prediction with the true new value when the model does not majorize $f$
  1. Gradient descent: $\x_{k+1}=\x_k-\alpha_k\g_k$; cheap ($O(n)$ beyond the gradient), so it scales to huge problems.
  2. Each step exactly minimizes the isotropic surrogate $f(\x_k)+\g_k^\top(\y-\x_k)+\frac1{2\alpha}\norm{\y-\x_k}^2$, which predicts the value $f(\x_k)-\frac\alpha2\norm{\g_k}^2$.
  3. On a quadratic the surrogate lies above $f$ exactly when $\alpha\le1/L$, which guarantees decrease; the step-size question is the heart of the method.

The surrogate $m_k(\y)=f(\x_k)+\g_k^\top(\y-\x_k)+\frac{1}{2\alpha}\norm{\y-\x_k}^2$ is minimized at…

$\y=\x_k-\g_k/\alpha$
Set $\grad m_k(\y)=\g_k+\frac1\alpha(\y-\x_k)$ to zero and solve carefully for $\y$.
$\y=\x_k$, because the model touches $f$ there
Touching is not the same as being lowest: the tangent term $\g_k^\top(\y-\x_k)$ is negative in some directions.
$\y=\x_k-\alpha\g_k$
Setting the gradient $\g_k+\frac1\alpha(\y-\x_k)$ to zero gives exactly the gradient step.

For $f(\x)=\frac12\x^\top A\x-\b^\top\x$ with eigenvalues of $A$ in $[1,20]$, the surrogate $m_k$ lies above $f$ for every $\x_k$ if and only if…

$\alpha\le1/20$
$m_k-f=\frac12(\y-\x_k)^\top(\frac1\alpha I-A)(\y-\x_k)\ge0$ for all $\y$ iff $\frac1\alpha\ge\lambda_{\max}=20$.
$\alpha\le1$
That only beats the smallest curvature. The model must be at least as curved as $f$ in its steepest direction.
$\alpha\lt 2/20$
$2/L$ will appear in Chapter 5.2 as the limit for convergence, but majorization is a stronger requirement.

Why does gradient descent scale to problems with $n=10^7$ variables while Newton's method does not?

Gradient descent needs fewer iterations
Usually it needs more iterations. The difference is in the cost of each one.
Each iteration costs $O(n)$ memory and work beyond the gradient, while Newton needs an $n\times n$ matrix and an $O(n^3)$ solve
$10^{14}$ matrix entries cannot even be stored; $10^7$ numbers can.
Gradient descent works without derivatives
It needs the gradient at every step.

On $f(x)=x^2$ ($L=2$), which fixed step lands exactly on the minimizer in one step from any $x_0$?

$\alpha=1$
Compute $x_1=x_0-\alpha\cdot2x_0$ with this $\alpha$: where do you land?
$\alpha=1/2$
$x_1=(1-2\alpha)x_0=0$. With $\alpha=1/L$ the surrogate coincides with $f$ itself.
$\alpha=2$
$x_1=(1-4)x_0=-3x_0$: further away than you started.

A fixed step on a quadratic bowl: when it works, and the best choice

On a convex quadratic, gradient descent with a fixed step converges from every start if and only if $0\lt\alpha\lt 2/L$, and the best fixed step $\alpha^\star=2/(m+L)$ shrinks the error by the factor $\frac{\kappa-1}{\kappa+1}$ per step.

This is the cleanest convergence analysis in the course and a standard exam question: derive the error recursion, decouple it in the eigenbasis, find the stability range and the optimal step.

A car's suspension on a bumpy road: stiff springs (large curvature) bounce violently if the damping step is too big; soft springs (small curvature) settle slowly whatever you do. One fixed setting has to serve both.

$f(\x)=\tfrac12\x^\top A\x-\b^\top\x$, with $A$ symmetric positive definite, eigenvalues $0\lt m=\lambda_1\le\dots\le\lambda_n=L$ and orthonormal eigenvectors $\u_1,\dots,\u_n$ (Part 0b). Then $\grad f(\x)=A\x-\b$, $\hess f=A$, the unique minimizer is $\x^\star=A^{-1}\b$, and $\kappa=L/m$.

The error is $\e_k=\x_k-\x^\star$, and the gap is $E(\x)=f(\x)-f(\x^\star)=\tfrac12(\x-\x^\star)^\top A(\x-\x^\star)$.

(The gap formula: expand $\tfrac12(\x-\x^\star)^\top A(\x-\x^\star)=\tfrac12\x^\top A\x-\x^{\star\top}A\x+\tfrac12\x^{\star\top}A\x^\star$, use $A\x^\star=\b$ so the middle term is $-\b^\top\x$, and note $f(\x^\star)=-\tfrac12\x^{\star\top}A\x^\star$.)

First, one variable: the whole story on one parabola

For $f(x)=\tfrac L2x^2$, $f'(x)=Lx$ and the update is $x_{k+1}=x_k-\alpha Lx_k=(1-\alpha L)x_k$, so $x_k=(1-\alpha L)^kx_0$. Everything depends on the single factor $1-\alpha L$:

Step sizeFactor $1-\alpha L$What the iterates do
$0\lt\alpha\lt1/L$in $(0,1)$crawl monotonically towards 0
$\alpha=1/L$$0$land on the minimizer in one step
$1/L\lt\alpha\lt2/L$in $(-1,0)$converge while ping-ponging across 0
$\alpha=2/L$$-1$bounce between $\pm x_0$ forever
$\alpha>2/L$below $-1$diverge, oscillating with growing amplitude

So $2/L$ is not a cautious technical bound: it is the exact threshold. In $n$ dimensions the same thing happens independently along each eigenvector, each with its own curvature $\lambda_i$. Before proving it, test your intuition.

A 2-D bowl has curvatures $m=1$ (along the valley) and $L=10$ (across it), so $2/L=0.2$. You run gradient descent with $\alpha=0.21$, just 5% above the limit, from a point far along the valley. What do you see?

It diverges immediately and in every direction
Along the valley the factor is $1-0.21\times1=0.79$: that component shrinks nicely. Only one direction is in trouble.
It converges, just more slowly, since 5% over is harmless
Across the valley the factor is $1-0.21\times10=-1.1$: that component flips sign and grows by 10% every step. Small growth compounds.
It first slides down the valley and seems to settle, then zig-zags across the valley with ever-growing swings and blows up
The flat component shrinks by $0.79$ per step while the steep one grows by $1.1$: after 10 steps $1.1^{10}\approx2.6$, after 50 steps $\approx117$.
Try it

This is the setting of the question ($\kappa=10$, $\alpha=2.1/L$). Press "Run 30" and watch the path and the gap chart below it. Then try $\alpha$ just below $2/L$ (1.95), exactly $1/L$, and the "Best fixed α" button. Rotate the bowl: nothing essential changes, because only the eigenvalues matter.

Step 1: the error recursion

Because $\b=A\x^\star$, the gradient is $\g_k=A\x_k-A\x^\star=A\e_k$. Subtract $\x^\star$ from both sides of $\x_{k+1}=\x_k-\alpha\g_k$: $$\e_{k+1}=\e_k-\alpha A\e_k=(I-\alpha A)\,\e_k,\qquad\text{so}\qquad \e_k=(I-\alpha A)^k\e_0 .$$

Step 2: decouple in the eigenbasis

Write the error in eigenvector coordinates: $\e_k=\sum_i z_{i,k}\u_i$, i.e. $\z_k=U^\top\e_k$ where $U$ has the eigenvectors as columns. Since $A\u_i=\lambda_i\u_i$, we get $(I-\alpha A)\u_i=(1-\alpha\lambda_i)\u_i$, so $$z_{i,k+1}=(1-\alpha\lambda_i)\,z_{i,k}\qquad\Longrightarrow\qquad z_{i,k}=(1-\alpha\lambda_i)^k z_{i,0}.$$ The $n$-dimensional problem has split into $n$ independent one-variable parabolas, one per eigen-direction, each with its own curvature $\lambda_i$. And because $U$ is orthogonal, $\norm{\e_k}=\norm{\z_k}$: lengths are the same in either coordinate system.

Try it

Left: the two contraction factors $|1-\alpha m|$ (green) and $|1-\alpha L|$ (blue) as you vary $\alpha$; the thick curve is the larger one, $\rho(\alpha)$. Right: each error component $z_{i,k}$ over 30 steps on a log scale (a hollow dot means the component is negative: it has flipped sign). Find the $\alpha$ where the thick curve is lowest, and explain why it sits where the two V's cross. Then increase $\kappa$ and watch the best $\rho$ creep towards 1.

With a constant step $\alpha$, gradient descent converges to $\x^\star$ from every starting point if and only if $0\lt\alpha\lt 2/L$. The error satisfies $$\norm{\e_k}\le\rho(\alpha)^k\norm{\e_0},\qquad \rho(\alpha)=\max_i|1-\alpha\lambda_i|=\max\big(|1-\alpha m|,\ |1-\alpha L|\big).$$ The best constant step is $\alpha^\star=\dfrac{2}{m+L}$, with $\rho(\alpha^\star)=\dfrac{L-m}{L+m}=\dfrac{\kappa-1}{\kappa+1}$.

Prove the theorem (this proof is examinable).

  1. From the decoupled recursion, $|z_{i,k}|=|1-\alpha\lambda_i|^k|z_{i,0}|\le\rho(\alpha)^k|z_{i,0}|$ for every $i$. Squaring and summing, $\norm{\z_k}^2\le\rho(\alpha)^{2k}\norm{\z_0}^2$, and $\norm{\e_k}=\norm{\z_k}$ gives $\norm{\e_k}\le\rho(\alpha)^k\norm{\e_0}$.

    Each component shrinks by its own factor; the worst factor bounds them all.

  2. Convergence from every start needs every factor below 1 in size: $|1-\alpha\lambda_i|\lt1$, i.e. $-1\lt1-\alpha\lambda_i\lt1$, i.e. $0\lt\alpha\lambda_i\lt2$ for all $i$. Since $\lambda_i>0$ this is $\alpha>0$ and $\alpha\lt2/\lambda_i$ for all $i$; the tightest is $\lambda_n=L$, giving $0\lt\alpha\lt2/L$.

    "If": then $\rho(\alpha)\lt1$ and step 1 gives convergence. "Only if": if $|1-\alpha\lambda_i|\ge1$ for some $i$, start with $\e_0=\u_i$; then $\e_k=(1-\alpha\lambda_i)^k\u_i$ never shrinks.

  3. $|1-\alpha\lambda|$ is a V-shaped function of $\lambda$ for fixed $\alpha$, so over $\lambda\in[m,L]$ its maximum is at an end: $\rho(\alpha)=\max(|1-\alpha m|,|1-\alpha L|)$.

    That is why only the extreme eigenvalues $m$ and $L$ matter, however many others there are.

  4. As $\alpha$ grows, $1-\alpha m$ decreases and $\alpha L-1$ increases. The larger of the two is smallest where they are equal, with opposite signs: $1-\alpha m=\alpha L-1$, so $\alpha^\star(m+L)=2$, $\alpha^\star=\dfrac2{m+L}$.

    Increasing $\alpha$ past this point helps the flat direction but hurts the steep one more, and vice versa. This balance is the crossing point of the two V's in the widget.

  5. $\rho(\alpha^\star)=1-\dfrac{2m}{m+L}=\dfrac{L-m}{L+m}=\dfrac{\kappa-1}{\kappa+1}$ (divide top and bottom by $m$).

    The rate depends only on the ratio $\kappa=L/m$: rescaling $f$ changes $\alpha^\star$ but not the speed.

How many iterations? The $O(\kappa\log(1/\eps))$ law

To guarantee $\norm{\e_k}\le\eps\norm{\e_0}$ it suffices that $\rho^k\le\eps$, i.e. $k\ge\dfrac{\ln(1/\eps)}{\ln(1/\rho)}$. With $\rho=\frac{\kappa-1}{\kappa+1}$, $$\ln\frac1\rho=\ln\Big(1+\frac{2}{\kappa-1}\Big)\ge\frac{2/(\kappa-1)}{1+2/(\kappa-1)}=\frac{2}{\kappa+1},$$ using $\ln(1+x)\ge\frac{x}{1+x}$ for $x>0$. So $k\ge\frac{\kappa+1}{2}\ln\frac1\eps$ steps are enough: the number of iterations grows linearly in $\kappa$ and only logarithmically in the accuracy. Doubling $\kappa$ roughly doubles the work. This $O(\kappa\log(1/\eps))$ count is one of the most quoted facts in the course.

The simpler choice $\alpha=1/L$ gives $\rho=\max(1-\frac1\kappa,0)=1-\frac1\kappa$, about twice as many steps: still $O(\kappa\log(1/\eps))$.

Why it is slow. The step must respect the steepest direction ($\alpha\lt2/L$), but progress along the flattest direction is then only a fraction $\alpha m\lt 2/\kappa$ of the remaining distance per step. One curvature limits the step; the other limits the progress.

The function value also decreases every step. Expanding the quadratic exactly, $$f(\x_k-\alpha\g_k)=f(\x_k)-\alpha\norm{\g_k}^2+\tfrac{\alpha^2}{2}\g_k^\top A\g_k\le f(\x_k)-\alpha\Big(1-\tfrac{\alpha L}{2}\Big)\norm{\g_k}^2,$$ using $\g_k^\top A\g_k\le L\norm{\g_k}^2$ (Rayleigh, Part 0b). For $0\lt\alpha\lt2/L$ the bracket is positive, so $f$ strictly decreases until $\g_k=\0$. And the gap, $E=\tfrac12\sum_i\lambda_iz_i^2$, shrinks by at least $\rho(\alpha)^2$ per step, since every $z_i$ shrinks by at least $\rho(\alpha)$.

Run gradient descent with $\alpha=1/L$ on $f(\x)=2x_1^2+\tfrac12x_2^2$ from $\x_0=(1,1)^\top$ (Tutorial 1, Problem 25). Find $\x_k$ and the ratio $\frac{f(\x_{k+1})}{f(\x_k)}$.

  1. $A=\mathrm{diag}(4,1)$: $L=4$, $m=1$, $\kappa=4$, $\x^\star=\0$, $f^\star=0$. $\grad f=(4x_1,\ x_2)^\top$ and $\alpha=1/4$.

    A diagonal $A$ means the coordinate axes are already the eigenvectors: no change of basis is needed.

  2. $x_{1,k+1}=(1-\tfrac14\cdot4)x_{1,k}=0$ and $x_{2,k+1}=(1-\tfrac14\cdot1)x_{2,k}=\tfrac34x_{2,k}$. So $\x_k=\big(0,(\tfrac34)^k\big)^\top$ for $k\ge1$; e.g. $\x_3=(0,\ 27/64)^\top$.

    The factors are $1-\alpha\lambda_i$: $0$ for the steep coordinate (killed in one step) and $3/4$ for the flat one.

  3. $f(\x_k)=\tfrac12(\tfrac34)^{2k}=\tfrac12(\tfrac9{16})^k$, so the ratio is $\tfrac9{16}=(1-\tfrac1\kappa)^2=0.5625$ every step.

    The gap is quadratic in the error, so it shrinks by the square of the error factor.

  4. With the best step $\alpha^\star=2/5$ instead, both factors equal $\pm\tfrac35=\pm\frac{\kappa-1}{\kappa+1}$, and the gap ratio is $\tfrac9{25}=0.36$ every step.

    $1-\tfrac25\cdot1=\tfrac35$ and $1-\tfrac25\cdot4=-\tfrac35$: balanced, exactly as the proof predicts.

Fit the line, with a fixed step

For the fit-the-line loss, $m=17-\sqrt{265}\approx0.7212$ and $L=17+\sqrt{265}\approx33.279$ (trace 34, determinant 24), so $\kappa\approx46.14$. Stability needs $\alpha\lt2/L\approx0.0601$. The best step is $\alpha^\star=2/34=1/17\approx0.0588$, with $\rho=\frac{L-m}{L+m}=\frac{\sqrt{265}}{17}\approx0.9576$. Shrinking the error by a factor $10^6$ is guaranteed after $\lceil\ln10^6/\ln(1/0.9576)\rceil=319$ steps (the estimate $\frac{\kappa+1}2\ln10^6\approx326$ agrees). A problem with just two unknowns needs hundreds of steps, purely because the valley is 46 times steeper across than along.

$f(\x)=\tfrac12\x^\top A\x$ with $A=\mathrm{diag}(1,10)$. (a) Fixed-step gradient descent converges from every start exactly when $0\lt\alpha\lt\ ?$ (b) What is the contraction factor $\rho(\alpha)$ for $\alpha=0.15$?

(a) $2/L$. (b) $\rho(\alpha)=\max(|1-\alpha m|,|1-\alpha L|)$ with $m=1$, $L=10$.

(a) $2/10=0.2$. (b) $|1-0.15|=0.85$ and $|1-1.5|=0.5$, so $\rho=0.85$: with this step the flat direction is the bottleneck.

Same $A=\mathrm{diag}(1,10)$. Find the best fixed step $\alpha^\star$ and its rate $\rho(\alpha^\star)$.

$\alpha^\star=2/(m+L)$ and $\rho=(\kappa-1)/(\kappa+1)$.

$\alpha^\star=2/11\approx0.1818$; $\rho=9/11\approx0.8182$. Check: $1-\tfrac2{11}=\tfrac9{11}$ and $1-\tfrac{20}{11}=-\tfrac9{11}$.

$f(x,y)=3x^2+2xy+3y^2$. Using the best fixed step, what is the smallest $k$ for which the bound $\rho^k$ guarantees $\norm{\e_k}\le10^{-6}\norm{\e_0}$?

Write $f=\tfrac12\x^\top A\x$: the diagonal of $A$ is $(6,6)$ and the off-diagonal is $2$. Find its eigenvalues, then $\rho$, then solve $\rho^k\le10^{-6}$.

$A=\begin{pmatrix}6&2\\2&6\end{pmatrix}$ has eigenvalues $6\pm2$, so $m=4$, $L=8$, $\kappa=2$, $\alpha^\star=1/6$, $\rho=1/3$. We need $3^{-k}\le10^{-6}$, i.e. $k\ge6\ln10/\ln3\approx12.58$, so $k=13$.

$A=\mathrm{diag}(1,10)$ again, and $\alpha=0.2$ exactly, from a generic start (both error components nonzero). What happens in the long run?

Compute both factors $1-\alpha\lambda_i$. What does a factor of exactly $-1$ do to a component?

Factors: $1-0.2=0.8$ (flat component dies out) and $1-2=-1$ (steep component flips sign but keeps its size). The iterates approach $\x^\star\pm z_{2,0}\u_2$ alternately: a 2-cycle, neither converging nor diverging. This is the boundary case $\alpha=2/L$.

  • Start every fixed-step analysis with $\e_{k+1}=(I-\alpha A)\e_k$ (it uses $\b=A\x^\star$)
  • Work in eigen-coordinates: each component is multiplied by $1-\alpha\lambda_i$
  • Remember only $m$ and $L$ matter for $\rho(\alpha)$
  • Quote the iteration count as $O(\kappa\log(1/\eps))$
  • Writing the stability condition as $\alpha\lt2/m$ (the steep direction, $L$, is the binding one)
  • Writing $\alpha\le2/L$: at $\alpha=2/L$ exactly it does not converge
  • Forgetting the absolute value in $|1-\alpha\lambda|$ (negative factors are fine if smaller than 1 in size)
  • Mixing up the rate for $\norm{\e_k}$ ($\rho$) with the rate for $f(\x_k)-f^\star$ ($\rho^2$)
  1. $\e_{k+1}=(I-\alpha A)\e_k$; in the eigenbasis each component is multiplied by $1-\alpha\lambda_i$ every step.
  2. Converges from every start iff $0\lt\alpha\lt2/L$; the rate is $\rho(\alpha)=\max(|1-\alpha m|,|1-\alpha L|)$, best at $\alpha^\star=\frac2{m+L}$ with $\rho=\frac{\kappa-1}{\kappa+1}$.
  3. Iterations needed grow like $\frac{\kappa+1}2\ln\frac1\eps$: linear in $\kappa$, which is why ill-conditioned valleys are slow.

A positive definite $A$ has eigenvalues $2, 5, 50$. Fixed-step gradient descent converges from every start iff…

$0\lt\alpha\lt1$
That is $2/\lambda$ for the smallest eigenvalue. Which direction blows up first as $\alpha$ grows?
$0\lt\alpha\lt0.04$
$2/L=2/50$. The middle eigenvalue 5 is irrelevant to stability.
$0\lt\alpha\le0.04$
At $\alpha=2/L$ the steep component is multiplied by $-1$ forever.

Same eigenvalues $2,5,50$. What is the best fixed step's rate $\rho(\alpha^\star)$?

$\frac{50-5}{50+5}$
The rate uses the two extreme eigenvalues only.
$1-\frac{2}{50}$
That is $\rho$ for $\alpha=1/L$, not the best step.
$\frac{48}{52}\approx0.923$
$\frac{L-m}{L+m}=\frac{50-2}{50+2}=\frac{\kappa-1}{\kappa+1}$ with $\kappa=25$.

Why is the best fixed step exactly where $1-\alpha m=-(1-\alpha L)$?

Because $\rho$ is the larger of a decreasing and an increasing function of $\alpha$, and that maximum is smallest where they cross
Move $\alpha$ either way from the crossing and one of the two factors gets bigger.
Because there the steep component is killed in one step
That happens at $\alpha=1/L$, which leaves the flat component slower than necessary.
Because there the iterates stop oscillating
At $\alpha^\star$ the steep factor is $-\rho\lt0$: that component still flips sign every step.

Problem A has $\kappa=100$, problem B has $\kappa=400$. With the best fixed step, to reach the same accuracy B needs about…

the same number of iterations
The count is roughly $\frac{\kappa+1}{2}\ln\frac1\eps$. Does it depend on $\kappa$?
4 times as many iterations
The count is linear in $\kappa$.
16 times as many iterations
That would be quadratic in $\kappa$; the bound is linear.

Exact line search: the zig-zag and the $\left(\frac{\kappa-1}{\kappa+1}\right)^2$ rate

Exact line search picks, at every step, the step size that makes $f$ as small as possible along the downhill ray; on a quadratic it is $\alpha_k=\dfrac{\g_k^\top\g_k}{\g_k^\top A\g_k}$.

It needs no knowledge of $m$ or $L$, yet it still zig-zags in narrow valleys, and the gap shrinks by at most $\left(\frac{\kappa-1}{\kappa+1}\right)^2$ per step. Both facts, and their proofs, are classic exam material.

Walking downhill in a straight line until the ground starts rising again, then turning. In a narrow ravine you keep hitting the opposite wall and turning back: lots of walking, little progress along the ravine.

Given a descent direction $\d_k$ (one with $\g_k^\top\d_k\lt0$), the line function is $\phi(\alpha)=f(\x_k+\alpha\d_k)$, and exact line search chooses $$\alpha_k=\argmin_{\alpha>0}\phi(\alpha).$$

By the chain rule, $\phi'(\alpha)=\grad f(\x_k+\alpha\d_k)^\top\d_k$: the slope of $f$ along the line. In particular $\phi'(0)=\g_k^\top\d_k\lt0$, so the line starts downhill.

If $\d_k=-\g_k$ and $\alpha_k$ comes from exact line search, then $\g_{k+1}^\top\g_k=0$.

Proof. $\alpha_k>0$ is an interior minimizer of the one-variable function $\phi$, so $\phi'(\alpha_k)=0$ (first-order condition, Part 3). But $\phi'(\alpha_k)=\grad f(\x_{k+1})^\top\d_k=-\g_{k+1}^\top\g_k$. $\square$

Geometrically: you stop exactly where the ray becomes tangent to a level set, and the new gradient is perpendicular to that level set (Part 0b), hence perpendicular to the ray you came along. So each new direction turns by exactly $90^\circ$.

For $f(\x)=\tfrac12\x^\top A\x-\b^\top\x$ with $A\succ0$ and $\d_k=-\g_k\ne\0$: $$\alpha_k=\frac{\g_k^\top\g_k}{\g_k^\top A\g_k}.$$

Proof. Expanding (the quadratic's Taylor formula is exact), $\phi(\alpha)=f(\x_k-\alpha\g_k)=f(\x_k)-\alpha\,\g_k^\top\g_k+\tfrac{\alpha^2}2\g_k^\top A\g_k$. This is an upward parabola in $\alpha$ (since $\g_k^\top A\g_k>0$), so its minimizer solves $\phi'(\alpha)=-\g_k^\top\g_k+\alpha\,\g_k^\top A\g_k=0$. $\square$

Two useful consequences. By the Rayleigh bounds $m\norm{\g}^2\le\g^\top A\g\le L\norm{\g}^2$, every exact step satisfies $\frac1L\le\alpha_k\le\frac1m$: sometimes it takes steps far longer than the fixed-step limit $2/L$, safely, because they are tailored to the current direction. And the cost per step is still one matrix–vector product: compute $A\g_k$ once, use it for $\alpha_k$, and update $\g_{k+1}=\g_k-\alpha_kA\g_k$.

Try it

Exact line search on a bowl with $\kappa=10$. Press "Step" a few times and read the angle between successive steps in the readout. Then raise $\kappa$ to 50 and click different starting points: the zig-zag gets tighter and the gap chart flattens. Find a start where it converges in one step (hint: along an axis of the ellipses) and one where the measured ratio sits right on the dashed worst-case line.

Two exact-line-search steps on $f(\x)=10x_1^2+x_2^2$ from $\x_0=(1,10)^\top$ (Tutorial 1, Problem 27). Check orthogonality and compare $\x_2$ with $\x_0$.

  1. $A=\mathrm{diag}(20,2)$, so $L=20$, $m=2$, $\kappa=10$, $\x^\star=\0$. $\g_0=(20\cdot1,\ 2\cdot10)^\top=(20,20)^\top$.

    $f=\tfrac12\x^\top A\x$ with $A=\mathrm{diag}(20,2)$ reproduces $10x_1^2+x_2^2$.

  2. $\g_0^\top\g_0=800$, $\g_0^\top A\g_0=20\cdot400+2\cdot400=8800$, so $\alpha_0=\frac1{11}$ and $\x_1=(1-\tfrac{20}{11},\ 10-\tfrac{20}{11})^\top=(-\tfrac9{11},\ \tfrac{90}{11})^\top$.

    The closed-form step $\g^\top\g/\g^\top A\g$.

  3. $\g_1=(20\cdot(-\tfrac9{11}),\ 2\cdot\tfrac{90}{11})^\top=\tfrac{180}{11}(-1,1)^\top$ and $\g_1^\top\g_0=\tfrac{180}{11}(-20+20)=0$.

    Orthogonality holds exactly, as the proposition promises.

  4. $\g_1^\top\g_1=2(\tfrac{180}{11})^2$, $\g_1^\top A\g_1=22(\tfrac{180}{11})^2$, so $\alpha_1=\tfrac1{11}$ and $\x_2=\x_1-\tfrac1{11}\g_1=(\tfrac{81}{121},\ \tfrac{810}{121})^\top=\left(\tfrac9{11}\right)^2\x_0$.

    After two steps the iterate points in the same direction as $\x_0$, scaled by $(\frac{\kappa-1}{\kappa+1})^2=(\frac9{11})^2$. The pattern repeats forever, and $f$ shrinks by $(\frac9{11})^2\approx0.669$ every single step: $f(\x_0)=110$, $f(\x_1)=\frac{8910}{121}\approx73.6$.

The convergence rate

Let $A\succ0$ have extreme eigenvalues $m$ and $L$, $\kappa=L/m$. For every $\y\ne\0$, $$\frac{(\y^\top\y)^2}{(\y^\top A\y)(\y^\top A^{-1}\y)}\ \ge\ \frac{4mL}{(m+L)^2}=\frac{4\kappa}{(1+\kappa)^2}.$$

You need the statement, not the proof (the book marks the proof as not examinable; it is in the box below). Recognize $\frac{4\kappa}{(1+\kappa)^2}$ on sight: $1-\frac{4\kappa}{(1+\kappa)^2}=\frac{(\kappa-1)^2}{(\kappa+1)^2}$.

Go deeper: a short proof of Kantorovich

Write $\y$ in eigen-coordinates, $\y=\sum_ic_i\u_i$, and let $\xi_i=c_i^2/\sum_jc_j^2$ (nonnegative weights summing to 1). Then the left side equals $1/(\Lambda M)$ with $\Lambda=\sum_i\xi_i\lambda_i$ and $M=\sum_i\xi_i/\lambda_i$.

For every $\lambda\in[m,L]$, $(\lambda-m)(\lambda-L)\le0$, i.e. $\lambda^2-(m+L)\lambda+mL\le0$; dividing by $\lambda$, $\lambda+\frac{mL}{\lambda}\le m+L$. Averaging with the weights $\xi_i$: $\Lambda+mL\,M\le m+L$.

By the AM–GM inequality, $2\sqrt{mL\,\Lambda M}\le\Lambda+mL\,M\le m+L$, so $\Lambda M\le\frac{(m+L)^2}{4mL}$, which is the claim. Equality needs all weight on $\lambda=m$ and $\lambda=L$, split evenly in the AM–GM sense: $\Lambda=mL\,M$. That is exactly the worst case for steepest descent below.

On $f(\x)=\tfrac12\x^\top A\x-\b^\top\x$ with $A\succ0$, gradient descent with exact line search satisfies, at every step, $$f(\x_{k+1})-f^\star\ \le\ \left(\frac{\kappa-1}{\kappa+1}\right)^2\big(f(\x_k)-f^\star\big).$$

Prove the theorem (examinable).

  1. Write $E_k=f(\x_k)-f^\star=\tfrac12\e_k^\top A\e_k$. Since $\g_k=A\e_k$, we have $\e_k=A^{-1}\g_k$ and $E_k=\tfrac12\g_k^\top A^{-1}\g_k$.

    We want everything in terms of $\g_k$, because $\alpha_k$ is expressed through $\g_k$.

  2. $\e_{k+1}=\e_k-\alpha_k\g_k$, so $E_{k+1}=\tfrac12(\e_k-\alpha_k\g_k)^\top A(\e_k-\alpha_k\g_k)=E_k-\alpha_k\g_k^\top A\e_k+\tfrac{\alpha_k^2}2\g_k^\top A\g_k=E_k-\alpha_k\g_k^\top\g_k+\tfrac{\alpha_k^2}2\g_k^\top A\g_k$.

    Expand the square; $A$ is symmetric so the two cross terms are equal; and $A\e_k=\g_k$.

  3. Substituting $\alpha_k=\frac{\g_k^\top\g_k}{\g_k^\top A\g_k}$: $E_k-E_{k+1}=\dfrac{(\g_k^\top\g_k)^2}{2\,\g_k^\top A\g_k}$.

    $-\frac{(\g^\top\g)^2}{\g^\top A\g}+\frac12\frac{(\g^\top\g)^2}{\g^\top A\g}=-\frac12\frac{(\g^\top\g)^2}{\g^\top A\g}$.

  4. Divide by $E_k=\tfrac12\g_k^\top A^{-1}\g_k$: $\dfrac{E_{k+1}}{E_k}=1-\dfrac{(\g_k^\top\g_k)^2}{(\g_k^\top A\g_k)(\g_k^\top A^{-1}\g_k)}$.

    This is [LD]'s Lemma 1: an exact formula for the one-step ratio. Everything now hinges on one fraction.

  5. Kantorovich with $\y=\g_k$ says the fraction is at least $\frac{4\kappa}{(1+\kappa)^2}$, so $\dfrac{E_{k+1}}{E_k}\le1-\dfrac{4\kappa}{(1+\kappa)^2}=\dfrac{(\kappa+1)^2-4\kappa}{(\kappa+1)^2}=\left(\dfrac{\kappa-1}{\kappa+1}\right)^2.$

    $(\kappa+1)^2-4\kappa=\kappa^2-2\kappa+1=(\kappa-1)^2$. $\square$

Reading the result

  • The bound is tight. Equality in Kantorovich happens when $\g_k$ has equal weight on the eigenvectors for $m$ and $L$, as in the worked example ($\g_0=(20,20)$). Then two steps reproduce the same situation scaled down, and the worst-case ratio is hit at every step. The book's lab measured per-step ratios $0.440$, $0.921$, $0.991$ for $\kappa=5,50,500$ against the bounds $0.444$, $0.923$, $0.992$: from a generic start the method settles into this worst-case zig-zag.
  • It can also be much faster. If $\g_0$ is an eigenvector, $\alpha_0=1/\lambda$ and the method lands on $\x^\star$ in one step. And if $\kappa=1$ (a round bowl) every start converges in one step.
  • Compare with the best fixed step. $\alpha^\star=\frac2{m+L}$ shrinks every eigen-component by at least $\frac{\kappa-1}{\kappa+1}$, so it also shrinks $E=\frac12\sum\lambda_iz_i^2$ by at least $\left(\frac{\kappa-1}{\kappa+1}\right)^2$: the same guarantee. Exact line search does not improve the worst case; its advantage is that it needs no knowledge of $m$ and $L$. Both need $O(\kappa\log(1/\eps))$ iterations.
  • Rates for $E$ versus rates for $\norm{\e}$. Since $\frac m2\norm{\e}^2\le E\le\frac L2\norm{\e}^2$, a rate $\rho^2$ for $E$ corresponds to a rate $\rho$ for $\norm{\e}$ (up to a constant factor $\sqrt\kappa$). Always say which quantity a rate refers to.
Try it

Each dot is a measured iteration count to shrink the gap $f-f^\star$ by $\eps$, on the 2-D bowl $\mathrm{diag}(1,\kappa)$; dashed curves are the guaranteed bounds. From the start $(\kappa,1)$ (equal gradient weights: the worst case) exact search sits on its bound. Switch to the start $(1,1)$: exact search becomes much faster, while the best fixed step does not change at all. Why? (In 2-D, both factors $1-\alpha^\star\lambda_i$ have size exactly $\rho$.)

Fit the line, solved by gradient descent

Our fit-the-line loss has $\kappa\approx46.14$, so the exact-line-search guarantee is $\left(\frac{\kappa-1}{\kappa+1}\right)^2=\frac{(L-m)^2}{(L+m)^2}=\frac{265}{289}\approx0.917$ per step: about 160 steps to shrink the gap by $10^6$. From the start $(w,c)=(0,3)$ the measured ratio is a steady $0.661$ (34 steps). From $(0,0)$ it is about $0.0008$ (2 steps): the first gradient $-\b=(-46,-20)$ points almost exactly along the steep eigenvector. Same problem, same method, wildly different speeds: the theorem is a worst-case guarantee.

Try it

Left: the data and the current line. Right: the loss map with the path. Compare the four step rules from the same start, counting steps until the readout reports the gap below $10^{-6}$ of the start. Then pick "α = 0.061": barely above $2/L\approx0.0601$, it slowly diverges. Drag the start around to find fast and slow starting lines for exact search.

$f(\x)=x_1^2+2x_2^2$, $\x_0=(2,1)^\top$. Do one step of steepest descent with exact line search: find $\alpha_0$ and $\x_1$.

$A=\mathrm{diag}(2,4)$, $\g_0=A\x_0$. Then $\alpha_0=\g_0^\top\g_0/\g_0^\top A\g_0$.

$\g_0=(4,4)^\top$, $\g_0^\top\g_0=32$, $\g_0^\top A\g_0=2\cdot16+4\cdot16=96$, so $\alpha_0=\tfrac13$ and $\x_1=(2-\tfrac43,\ 1-\tfrac43)=(\tfrac23,-\tfrac13)^\top$. Check: $\g_1=(\tfrac43,-\tfrac43)^\top\perp\g_0$, and $f$ drops from 6 to $\tfrac23$, a ratio of $\tfrac19=\left(\frac{2-1}{2+1}\right)^2$: the worst case, because $\g_0$ has equal components.

Same $f(\x)=x_1^2+2x_2^2$, now from $\x_0=(1,1)^\top$. Compute the ratio $f(\x_1)/f(\x_0)$ after one exact-line-search step, and compare it with the bound $1/9$.

$\g_0=(2,4)^\top$. You can also use the exact one-step formula $E_1/E_0=1-\frac{(\g^\top\g)^2}{(\g^\top A\g)(\g^\top A^{-1}\g)}$.

$\g_0^\top\g_0=20$, $\g_0^\top A\g_0=8+64=72$, so $\alpha_0=\tfrac5{18}$ and $\x_1=(1-\tfrac{10}{18},\ 1-\tfrac{20}{18})=(\tfrac49,-\tfrac19)^\top$. $f(\x_0)=3$, $f(\x_1)=\tfrac{16}{81}+\tfrac{2}{81}=\tfrac29$, ratio $\tfrac2{27}\approx0.074\lt\tfrac19$. Via the formula: $\g^\top A^{-1}\g=\tfrac42+\tfrac{16}4=6$, so $1-\tfrac{400}{72\cdot6}=\tfrac{2}{27}$.

For the fit-the-line loss ($m=17-\sqrt{265}$, $L=17+\sqrt{265}$), give the exact-line-search guaranteed ratio $\left(\frac{L-m}{L+m}\right)^2$ as a number, and the smallest $k$ for which it guarantees $L(\z_k)-L^\star\le10^{-6}\big(L(\z_0)-L^\star\big)$.

$L-m=2\sqrt{265}$ and $L+m=34$. Then solve $\rho^{2k}\le10^{-6}$ with logarithms.

$\left(\frac{2\sqrt{265}}{34}\right)^2=\frac{4\cdot265}{1156}=\frac{265}{289}\approx0.91696$. Then $k\ge\frac{\ln10^6}{\ln(289/265)}=\frac{13.816}{0.08670}\approx159.4$, so $k=160$.

$A$ has eigenvalues $1$ and $100$. Can exact line search ever choose a step $\alpha_k$ larger than $2/L=0.02$ (the fixed-step stability limit)?

Use the Rayleigh bounds on $\g^\top A\g$ to find the range of $\alpha_k=\g^\top\g/\g^\top A\g$.

Yes. $\alpha_k\in[1/L,1/m]=[0.01,1]$; for example if $\g_k$ is the eigenvector for $\lambda=1$, then $\alpha_k=1$, fifty times the fixed-step limit, and it lands exactly on $\x^\star$. A long step is unsafe only when it is used in every direction; exact search tailors it to the current one.

  • Use $\alpha_k=\g_k^\top\g_k/\g_k^\top A\g_k$ on quadratics, and check $\g_{k+1}^\top\g_k=0$ as a sanity test
  • In proofs, rewrite the gap as $E=\frac12\g^\top A^{-1}\g$ before using Kantorovich
  • State which quantity a rate applies to: $f-f^\star$ (squared rate) or $\norm{\x-\x^\star}$
  • Explain zig-zag with orthogonality plus a long narrow valley
  • Thinking exact line search beats the best fixed step's worst-case rate (it doesn't; it only removes the need to know $m$, $L$)
  • Writing $\alpha_k=\g^\top A\g/\g^\top\g$ (upside down)
  • Claiming the rate is achieved from every start (it's an upper bound; some starts converge in one step)
  • Expecting a closed-form exact step for non-quadratic $f$ (then the line search needs its own iteration)
  1. Exact line search on a quadratic: $\alpha_k=\frac{\g_k^\top\g_k}{\g_k^\top A\g_k}\in[\frac1L,\frac1m]$, and successive gradients are orthogonal, which causes zig-zag in narrow valleys.
  2. $E_{k+1}/E_k=1-\frac{(\g^\top\g)^2}{(\g^\top A\g)(\g^\top A^{-1}\g)}\le\left(\frac{\kappa-1}{\kappa+1}\right)^2$ by Kantorovich; the bound is attained when the gradient weights the two extreme eigenvectors equally.
  3. Exact search and the best fixed step share the guarantee $\left(\frac{\kappa-1}{\kappa+1}\right)^2$ on $f-f^\star$ and the $O(\kappa\log(1/\eps))$ iteration count.

Why are successive steepest-descent directions perpendicular under exact line search?

Because $A$ is symmetric, so its eigenvectors are perpendicular
The orthogonality holds for any differentiable $f$, quadratic or not. What condition does the exact step satisfy?
Because $\phi'(\alpha_k)=\grad f(\x_{k+1})^\top\d_k=0$ at the minimizing step, and $\d_k=-\g_k$
The first-order condition for the one-variable minimization is exactly the orthogonality.
Because the step size is chosen as $1/L$
A fixed $1/L$ step does not make successive gradients orthogonal in general.

$f(x,y)=3x^2+2xy+3y^2$ (eigenvalues of $A$: 4 and 8). The guaranteed per-step reduction of $f-f^\star$ under exact line search is…

$\frac13$
That's $\frac{\kappa-1}{\kappa+1}$, the rate for the error norm with the best fixed step. Which power applies to the function gap?
$\frac19$
$\kappa=2$, so $\left(\frac{\kappa-1}{\kappa+1}\right)^2=\frac19$: each step gains about one decimal digit.
$\frac12$
That's $m/L$. The rate is built from $\kappa$ through $\frac{\kappa-1}{\kappa+1}$.

In the proof of the exact-line-search rate, what is the role of the Kantorovich inequality?

It shows $\alpha_k\le1/m$
That follows from the Rayleigh bounds. Kantorovich is used on a different fraction.
It proves successive gradients are orthogonal
Orthogonality comes from $\phi'(\alpha_k)=0$.
It lower-bounds $\frac{(\g^\top\g)^2}{(\g^\top A\g)(\g^\top A^{-1}\g)}$, the fraction of the gap removed in one step
Exactly: $E_{k+1}/E_k=1-$ that fraction $\le1-\frac{4\kappa}{(1+\kappa)^2}$.

On a bowl with $\kappa=1000$, exact line search from some start converges in a single step. This…

contradicts the theorem, which says the rate is $\left(\frac{999}{1001}\right)^2$
The theorem gives an upper bound on the ratio, not its exact value.
is possible: if $\g_0$ is an eigenvector of $A$, then $\x_1=\x^\star$
Then $\alpha_0=1/\lambda$ and $\e_1=\e_0-\frac1\lambda\lambda\e_0=\0$. The worst case needs both extreme eigenvectors present.
is impossible unless $\kappa=1$
Try a start on one of the ellipse's axes in the stepper.

$A\succ0$ has eigenvalues 1 and 100. Which pair of facts is right?

Level ellipses have axis ratio 100; fixed-step GD needs $\alpha\lt0.02$
Half-axes are $1/\sqrt{\lambda_i}$, so the axis ratio is $\sqrt\kappa$.
Level ellipses have axis ratio 10; fixed-step GD needs $\alpha\lt0.02$
$\sqrt{100}=10$, and $2/L=2/100$.
Level ellipses have axis ratio 10; fixed-step GD needs $\alpha\lt2$
$2/m$ would let the steep direction blow up. Which eigenvalue limits the step?

For $f(\x)=\frac12\x^\top A\x-\b^\top\x$ with $A\succ0$, a point is left unchanged by a gradient step ($\x=\x-\alpha\grad f(\x)$ with $\alpha>0$) exactly when…

$\x$ is a local maximum
A convex quadratic with $A\succ0$ has no local maxima.
$\alpha=2/L$
The fixed point condition does not depend on which positive $\alpha$ you use.
$A\x=\b$, i.e. $\x=\x^\star$, the unique global minimizer
Fixed points are stationary points; for a strictly convex quadratic the only one is $A^{-1}\b$.

Fixed-step GD cycles forever on $|x|$ but converges on every convex quadratic when $0\lt\alpha\lt2/L$. The key difference is…

On a quadratic, $\g_k=A\e_k$ shrinks in proportion to the error, so steps shrink as you approach; $|x|$ has slope $\pm1$ right up to the minimum
The proportionality is what produces the factor $1-\alpha\lambda_i$ in each eigen-direction.
$|x|$ is not convex
It is convex; it is not differentiable at 0.
On $|x|$ the step is too small
A smaller fixed step still cycles, just with smaller swings.

On a quadratic with $\kappa=50$, which statement about exact line search versus the best fixed step $\alpha^\star=2/(m+L)$ is correct?

Exact line search has a strictly better worst-case rate
Both shrink $f-f^\star$ by at least $\left(\frac{\kappa-1}{\kappa+1}\right)^2$ per step.
They have the same worst-case guarantee, but exact search needs no knowledge of $m$ and $L$
That is its practical advantage; it also adapts to lucky starts.
The best fixed step zig-zags while exact search does not
Exact search is the one that turns $90^\circ$ every step.

Why is the exact step $\alpha_k$ on a quadratic unique and positive whenever $\g_k\ne\0$?

Because $f$ is bounded below
Boundedness alone doesn't pin down a unique minimizer along the line.
Because $\alpha_k$ is always $1/L$
It varies between $1/L$ and $1/m$.
$\phi(\alpha)$ is an upward parabola (coefficient $\frac12\g_k^\top A\g_k>0$, strictly convex) with $\phi'(0)=-\norm{\g_k}^2\lt0$
A strictly convex 1-D function that starts downhill has a unique minimizer at some $\alpha>0$.

Lecture 7 (27 Aug). Part 5 chose the step size exactly, which only works cheaply on quadratics. This part replaces "the best step" by "a good enough step": a cheap test that rejects steps that are too long and steps that are too short. On the way we meet the single most useful inequality of the course, the descent lemma, and we prove that backtracking line search always finishes with a step that is safely large.

You need: Part 0b (directional derivative $\grad f^\top\d$, eigenvalues, the integral form of Taylor's theorem), Part 0a (Cauchy–Schwarz), and Part 5 (gradient descent, exact line search on quadratics, the condition $0\lt\alpha\lt2/L$).

Why the exact step is the wrong tool, and what can go wrong instead

Once you have picked a downhill direction, choosing how far to walk is a one-variable problem, and you don't need to solve it exactly: you only need a step that is neither too long nor too short.

Every practical optimizer (gradient descent in machine learning libraries, quasi-Newton codes, nonlinear CG) uses an inexact line search. Exam questions ask you to run one by hand and to explain why its two tests are needed.

Parking a car: you don't compute the exact centimetre where the car should stop. You stop once you're inside the bay, not short of it and not into the wall.

The line-search template

Every method in Parts 5–10 has the same shape. At the current point $\x_k$:

  1. choose a direction $\d_k$ that goes downhill (steepest descent uses $\d_k=-\g_k$, Newton uses something else);
  2. choose a step size $\alpha_k\gt0$;
  3. move: $\x_{k+1}=\x_k+\alpha_k\d_k$.

This part is entirely about step 2. Once $\x_k$ and $\d_k$ are fixed, the only unknown is the number $\alpha$, so we look at $f$ along the ray only.

Notation for this part
$\g_k$
the gradient $\grad f(\x_k)$ at the current point
$\d_k$
the search direction; a descent direction means $\g_k^\top\d_k\lt0$
$\phi(\alpha)$
"phi of alpha": $f(\x_k+\alpha\d_k)$, the height of the landscape after a step of size $\alpha$ along $\d_k$
$\alpha^\star$
an exact line-search step, a minimizer of $\phi$ over $\alpha\gt0$

For fixed $\x_k$ and $\d_k$, let $\phi(\alpha)=f(\x_k+\alpha\d_k)$ for $\alpha\ge0$. Then $$\phi(0)=f(\x_k),\qquad \phi'(\alpha)=\grad f(\x_k+\alpha\d_k)^\top\d_k,\qquad \phi'(0)=\g_k^\top\d_k\lt0.$$

The formula for $\phi'$ is the chain rule from Part 0b: moving $\alpha$ by a little moves the point along $\d_k$, and the slope of $f$ along $\d_k$ is $\grad f^\top\d_k$. So $\phi'(0)$ is the directional derivative at the start. It is negative exactly because $\d_k$ is a descent direction: the graph of $\phi$ starts by going down.

For steepest descent, $\d_k=-\g_k$ and $\phi'(0)=-\g_k^\top\g_k=-\norm{\g_k}^2$.

Try it

Drag the black point $\x_k$ on the map and turn the direction $\d$ (angle measured from $-\grad f$). The small plot shows the slice $\phi(\alpha)$ along the purple ray, its tangent at $\alpha=0$ (slope $\phi'(0)$), and the exact minimizer along the ray. Turn $\d$ past $90°$ from $-\grad f$ and watch $\phi'(0)$ change sign. On Rosenbrock's function, notice how $\phi$ can have a sharp narrow dip followed by a huge hill.

Why not just minimize φ exactly?

In Part 5 we did: for a quadratic $f(\x)=\frac12\x^\top A\x-\b^\top\x$, $\phi$ is a parabola in $\alpha$, and its minimizer has the closed form $\alpha_k=\g_k^\top\g_k/\g_k^\top A\g_k$. For anything else, $\phi'(\alpha)=0$ is a nonlinear equation.

Take $f(x_1,x_2)=x_1^4+2x_2^4$ at $\x_0=(1,1)^\top$ with $\d_0=-\g_0=-(4,8)^\top$. Then $\phi(\alpha)=(1-4\alpha)^4+2(1-8\alpha)^4$ and $$\phi'(\alpha)=-16(1-4\alpha)^3-64(1-8\alpha)^3=0$$ is a cubic with no clean root. A numerical solver finds $\alpha^\star\approx0.1549$. For a non-polynomial $f$ the equation is transcendental. Either way you must run an inner iterative solver (Newton on $\phi'$, or golden-section search) at every outer iteration, paying many function evaluations for one step. And the payoff is small: the next direction will change anyway, so a perfect step along this one is wasted precision.

So we ask for less: a step that is good enough to make the convergence proofs of Part 7 work, and that is cheap to find.

Is "f goes down" good enough?

The most naive rule is: accept any $\alpha$ with $\phi(\alpha)\lt\phi(0)$. It fails, in two opposite ways. Both are visible on the simplest function there is.

Minimize $f(x)=x^2$ (minimizer $x^\star=0$) by steepest descent from $x_0=1$, so $x_{k+1}=x_k-\alpha_k\cdot2x_k=(1-2\alpha_k)x_k$. Use the shrinking steps $\alpha_k=2^{-(k+2)}$, i.e. $\tfrac14,\tfrac18,\tfrac1{16},\dots$ Every step makes $f$ strictly smaller. Where do the iterates go?

To $0$, because $f$ decreases at every step
Decreasing is not the same as decreasing to the minimum. How much does each step still move $x$ once $\alpha_k$ is tiny?
To about $0.289$, a point that is not a minimizer
$x_k=\prod_j(1-2^{-(j+1)})$, and because the steps add up to a finite total, the product stops shrinking at $\approx0.2888$, where $f'\approx0.578\neq0$.
They diverge
Each factor $1-2\alpha_k$ is between 0 and 1, so $|x_k|$ can only shrink.
Try it

Choose a step schedule and press "Run". "Short" steps $\alpha_k=2^{-(k+2)}$ stall on the right. "Long" steps $\alpha_k=1-2^{-(k+2)}$ overshoot the minimum each time and stall at the same distance, bouncing from side to side. The table shows which steps pass the two tests of Chapter 6.3 (Armijo rejects the long ones, the curvature test rejects the short ones). Change $c_1$ and $c_2$ and see when each schedule starts failing.

The two failure modes

For $f(x)=x^2$ and $\d_k=-2x_k$, the exact step is $\alpha=\tfrac12$ (it lands on $0$). The schedules in the widget are mirror images around it:

  • Steps too short ($\alpha_k=\tfrac14,\tfrac18,\dots$). Each step decreases $f$, but the steps are summable ($\sum_k\alpha_k=\tfrac12\lt\infty$), so the total distance the method can ever travel is finite. The iterates run out of steam at $x_\infty\approx0.2888$.
  • Steps too long ($\alpha_k=\tfrac34,\tfrac78,\dots$). Then $1-2\alpha_k=-(1-2^{-(k+1)})$: the iterate jumps over the minimum and lands almost as far away on the other side. $|x_k|$ is exactly the same as before, so again it stalls at distance $0.2888$, while $f$ decreases each time by a vanishing amount.

[NW] Fig. 3.2 shows the same disease with function values $5/k\to0$ while the true minimum is $-1$. The cure is two tests: one that demands a decrease proportional to the step (this kills overly long steps), and one that refuses steps that stop while $\phi$ is still steeply decreasing (this kills overly short ones). Chapters 6.3 and 6.4 build them. Chapter 6.2 builds the tool that proves they work.

Let $f(x_1,x_2)=e^{x_1}+x_2^2$, $\x_k=(0,1)^\top$, $\d_k=-\g_k$. Write $\phi(\alpha)$, find $\phi'(0)$, compute the exact step, and compare it with the simple guess $\alpha=0.5$.

  1. $\grad f=(e^{x_1},\,2x_2)^\top$, so $\g_k=(1,2)^\top$ and $\d_k=(-1,-2)^\top$.

    Steepest descent: the direction is minus the gradient.

  2. $\x_k+\alpha\d_k=(-\alpha,\,1-2\alpha)^\top$, so $\phi(\alpha)=e^{-\alpha}+(1-2\alpha)^2$, with $\phi(0)=2$.

    Substitute the moving point into $f$: $\phi$ is an ordinary function of one variable.

  3. $\phi'(\alpha)=-e^{-\alpha}-4(1-2\alpha)$, so $\phi'(0)=-1-4=-5=-\norm{\g_k}^2$.

    Two checks agree: differentiating $\phi$ directly, and the formula $\phi'(0)=\g_k^\top\d_k$.

  4. Exact step: $\phi'(\alpha)=0\iff 8\alpha=4+e^{-\alpha}$. This mixes $\alpha$ and $e^{-\alpha}$, so it has no closed-form solution. Iterating $\alpha\leftarrow(4+e^{-\alpha})/8$ from $0.5$ gives $\alpha^\star\approx0.5706$, with $\phi(\alpha^\star)\approx0.5851$.

    A transcendental equation: we needed an inner iterative solver just to choose one step.

  5. The guess $\alpha=0.5$ gives $\phi(0.5)=e^{-0.5}+0\approx0.6065$. It achieves $(2-0.6065)/(2-0.5851)\approx98.5\%$ of the best possible decrease along this ray, at the cost of one function evaluation.

    This is the whole case for inexact line search: almost all of the benefit for almost none of the cost.

Let $f(x_1,x_2)=x_1^2x_2+x_2^3$, $\x_k=(1,1)^\top$, $\d_k=(1,-1)^\top$. Compute $\phi'(0)$ and $\phi(0.5)$, where $\phi(\alpha)=f(\x_k+\alpha\d_k)$.

$\grad f=(2x_1x_2,\ x_1^2+3x_2^2)^\top$. For $\phi(0.5)$, the point is $(1.5,\,0.5)$.

$\g_k=(2,4)^\top$, so $\phi'(0)=\g_k^\top\d_k=2-4=-2$: a descent direction. At $\alpha=0.5$ the point is $(1.5,0.5)$ and $\phi(0.5)=(2.25)(0.5)+0.125=1.25$, below $\phi(0)=2$.

For $f(x)=x^2$, $x_0=1$, $x_{k+1}=(1-2\alpha_k)x_k$: compute $x_3$ for the short schedule $\alpha_k=2^{-(k+2)}$ and for the long schedule $\alpha_k=1-2^{-(k+2)}$.

Short: the factors are $1-\tfrac12,\ 1-\tfrac14,\ 1-\tfrac18$. Long: each factor is the negative of the corresponding short one.

Short: $x_3=\tfrac12\cdot\tfrac34\cdot\tfrac78=\tfrac{21}{64}=0.328125$. Long: $x_3=(-\tfrac12)(-\tfrac34)(-\tfrac78)=-0.328125$. Same distance from $0$, opposite side.

For $f(x)=x^2$ at any $x_k\neq0$ with $\d_k=-2x_k$, the step $\alpha$ decreases $f$ exactly when $0\lt\alpha\lt\bar a$. Find $\bar a$ and the exact line-search step.

$f(x_{k+1})=(1-2\alpha)^2x_k^2$. When is $|1-2\alpha|\lt1$?

$|1-2\alpha|\lt1\iff0\lt\alpha\lt1$, so $\bar a=1$. The exact step makes $1-2\alpha=0$: $\alpha^\star=\tfrac12$. The decreasing range is symmetric around the exact step, which is why the short and long schedules mirror each other.

For $f(x_1,x_2)=x_1^4+2x_2^4$ at $(1,1)$ with $\d_0=-\g_0$, compute $\phi'(0)$ directly from $\phi(\alpha)=(1-4\alpha)^4+2(1-8\alpha)^4$, and check it equals $-\norm{\g_0}^2$.

Chain rule: $\frac{d}{d\alpha}(1-4\alpha)^4=-16(1-4\alpha)^3$.

$\phi'(\alpha)=-16(1-4\alpha)^3-64(1-8\alpha)^3$, so $\phi'(0)=-80$. And $\g_0=(4,8)^\top$, $\norm{\g_0}^2=16+64=80$.

  • Reduce step-size questions to the one-variable function $\phi(\alpha)$
  • Check $\phi'(0)=\g_k^\top\d_k\lt0$ before any line search: no descent direction, no line search
  • Use the closed-form exact step only for quadratics
  • Think of a step as needing two guarantees: not too long, not too short
  • Believing "$f$ decreases every iteration" implies convergence to a minimizer (the $x^2$ schedules refute it)
  • Spending an inner solver on the exact step for a general $f$
  • Forgetting the factor $\d_k$ in $\phi'(\alpha)=\grad f(\x_k+\alpha\d_k)^\top\d_k$
  1. A line search chooses $\alpha$ by looking only at $\phi(\alpha)=f(\x_k+\alpha\d_k)$, which starts with slope $\phi'(0)=\g_k^\top\d_k\lt0$.
  2. Exact minimization of $\phi$ needs its own iterative solver except for quadratics, and buys little.
  3. "Any decrease" is not enough: steps that are too short (summable) or too long (overshooting) can both stall at a non-stationary point.

At $\x_k$, $\g_k=(3,-1)^\top$ and $\d_k=(-1,-2)^\top$. What is $\phi'(0)$, and is $\d_k$ usable for a line search?

$\phi'(0)=5$, usable
Recompute the dot product; watch the signs of both terms.
$\phi'(0)=1$, not usable since it is positive
Check the sign of the product $(-1)\cdot(-2)$: it multiplies $g_2=-1$ by $d_2=-2$.
$\phi'(0)=-1$, usable
$3(-1)+(-1)(-2)=-3+2=-1\lt0$: a descent direction, so small steps decrease $f$.

Why is exact line search standard for quadratics but not for general $f$?

For a quadratic, $\phi$ is a parabola whose minimizer has a closed form; in general $\phi'(\alpha)=0$ needs an iterative solver at every iteration
Exactly: $\alpha_k=\g_k^\top\g_k/\g_k^\top A\g_k$ for quadratics, an inner loop otherwise.
Because exact line search diverges on non-quadratics
It doesn't diverge; the issue is cost, not stability.
Because non-quadratic $\phi$ has no minimizer
It usually has one; the issue is computing it.

Gradient descent on $x^2$ with $\alpha_k=2^{-(k+2)}$ decreases $f$ every step but stalls. What property of the step sequence causes this?

The steps are larger than $1/L$
Here $L=2$, so $1/L=\tfrac12$, and every step is at most $\tfrac14$.
The steps are summable, $\sum_k\alpha_k\lt\infty$
The product $\prod(1-2\alpha_k)$ then converges to a positive number: the method can only travel a finite total distance.
The direction is not a descent direction
$\d_k=-2x_k$ is steepest descent, always downhill.

For $f=x^2$, the "long" schedule $\alpha_k=1-2^{-(k+2)}$ also stalls. Compared with the "short" schedule, its iterates…

converge to $0$, because long steps cover more ground
Compute $|1-2\alpha_k|$ for both schedules: the two are equal.
diverge
$|1-2\alpha_k|\lt1$ for every $k$, so $|x_k|$ shrinks.
have exactly the same $|x_k|$ but alternate in sign
$1-2\alpha_k=-(1-2^{-(k+1)})$, the negative of the short-schedule factor.

Lipschitz gradients and the descent lemma

If the gradient of $f$ cannot change faster than a rate $L$, then $f$ lies below a parabola of curvature $L$ that touches it at any point you choose. That parabola tells you exactly which steps are safe.

The descent lemma is the workhorse of every convergence proof from here to Part 8: backtracking (this part), constant steps and Zoutendijk (Part 7), and the $O(1/k)$ rates (Part 8). Its proof is examinable.

A speed limit on a road: if you know nobody can drive faster than 100 km/h, you can bound how far a car got from where it was seen, without tracking it. $L$ is a speed limit on how fast the slope can change.

Lipschitz continuity

A function $h$ is Lipschitz continuous with constant $L$ if $\norm{h(\x)-h(\y)}\le L\norm{\x-\y}$ for all $\x,\y$. In words: the output can't change faster than $L$ times the input change. In one variable: every chord of the graph has slope between $-L$ and $L$.

  • $h(x)=|x|$ is Lipschitz with $L=1$ (kink and all).
  • $h(x)=x^2$ is not Lipschitz on $\R$: its chords get arbitrarily steep. On $[-3,3]$ it is, with $L=6$.
  • $h(x)=\sqrt{|x|}$ is continuous but not Lipschitz near $0$: chords from $0$ have slope $1/\sqrt{|x|}\to\infty$.

For optimization, the useful thing to bound is not $f$ itself but its gradient.

$f\in C^1(\R^n)$ is $L$-smooth if its gradient is $L$-Lipschitz: $$\norm{\grad f(\x)-\grad f(\y)}\le L\norm{\x-\y}\qquad\text{for all }\x,\y\in\R^n.$$ In Nesterov's notation (Part 8), a convex $L$-smooth function belongs to the class $\mathcal F_L^{1,1}$.

If $f\in C^2$, then $f$ is $L$-smooth if and only if every eigenvalue of $\hess f(\x)$ lies in $[-L,L]$, for every $\x$. When $f$ is convex (all eigenvalues $\ge0$), this is simply $\hess f(\x)\preceq LI$, i.e. $\lambda_{\max}(\hess f(\x))\le L$ everywhere.

The smallest valid $L$ is the largest curvature (in absolute value) anywhere. Examples:

  • Quadratics $f=\frac12\x^\top A\x-\b^\top\x$ with $A\succ0$: $\hess f=A$, so $L=\lambda_{\max}(A)$. This is why the book calls the largest eigenvalue $L$. Our fit-the-line loss has $L\approx33.28$.
  • $f(x)=\ln(1+e^x)$ (the logistic loss): $f''(x)=\sigma(x)(1-\sigma(x))$ with $\sigma(x)=1/(1+e^{-x})$, and this is at most $\tfrac14$. So $L=\tfrac14$.
  • $f(x)=x^4$: $f''(x)=12x^2$ is unbounded, so $x^4$ is not $L$-smooth on $\R$ for any $L$. (It is on a bounded interval.) Tutorial 1 Q23 shows the consequence: for every fixed step size there are starting points from which gradient descent on $x^4$ diverges.
Careful: for a non-convex $f$, "$\hess f\preceq LI$" alone is not the same as $L$-smoothness: $f(x)=-x^2$ has $f''=-2\le0$, yet its gradient $-2x$ is only $2$-Lipschitz, not $0$-Lipschitz. The descent lemma below only uses the upper bound on curvature, which is why some texts state the test one-sidedly.

If $f$ is $L$-smooth, then for all $\x,\y\in\R^n$ $$f(\y)\le f(\x)+\grad f(\x)^\top(\y-\x)+\frac L2\norm{\y-\x}^2.$$ No convexity is assumed.

Reading it. The first two terms on the right are the tangent plane at $\x$. The lemma says: $f$ can rise above its tangent by at most $\frac L2\norm{\y-\x}^2$. So $f$ is trapped below a parabola (a bowl in $n$ dimensions) with curvature $L$ that touches $f$ at $\x$. It is an upper bound only: $f$ may fall far below it, especially when $f$ is not convex.

Prove the descent lemma. (Examinable: Book Ch. 11, Tutorial 1 Q22(a).)

  1. Walk along the segment: $\z(t)=\x+t(\y-\x)$, $t\in[0,1]$. By the integral form of Taylor's theorem (Part 0b), $$f(\y)-f(\x)=\int_0^1\grad f(\z(t))^\top(\y-\x)\,dt.$$

    This turns a statement about two points into a one-variable integral, exactly the trick of Part 0b. It is exact, needs only $f\in C^1$, and works for any distance between $\x$ and $\y$.

  2. Subtract $\grad f(\x)^\top(\y-\x)=\int_0^1\grad f(\x)^\top(\y-\x)\,dt$ from both sides: $$f(\y)-f(\x)-\grad f(\x)^\top(\y-\x)=\int_0^1\big(\grad f(\z(t))-\grad f(\x)\big)^\top(\y-\x)\,dt.$$

    The left side is "how far $f$ rises above its tangent"; the integrand compares the gradient along the way with the gradient at the start, which is what the Lipschitz condition controls.

  3. Cauchy–Schwarz, then $L$-smoothness with $\z(t)-\x=t(\y-\x)$: $$\big(\grad f(\z(t))-\grad f(\x)\big)^\top(\y-\x)\le\norm{\grad f(\z(t))-\grad f(\x)}\,\norm{\y-\x}\le Lt\norm{\y-\x}^2.$$

    Cauchy–Schwarz (Part 0a) converts the dot product into a product of lengths; the Lipschitz bound replaces the gradient difference by $L$ times the distance travelled, $t\norm{\y-\x}$.

  4. Integrate: $\int_0^1Lt\norm{\y-\x}^2\,dt=\frac L2\norm{\y-\x}^2$. Hence $f(\y)\le f(\x)+\grad f(\x)^\top(\y-\x)+\frac L2\norm{\y-\x}^2$. $\blacksquare$

    The $\frac12$ comes from $\int_0^1t\,dt$. Nowhere did we use convexity: only $C^1$, Cauchy–Schwarz and the Lipschitz gradient.

If $f$ is $L$-smooth and $\x^+=\x-\alpha\grad f(\x)$, then $$f(\x^+)\le f(\x)-\alpha\Big(1-\frac{L\alpha}2\Big)\norm{\grad f(\x)}^2.$$ The guaranteed decrease is positive for $0\lt\alpha\lt2/L$ and largest at $\alpha=1/L$, where $$f(\x^+)\le f(\x)-\frac1{2L}\norm{\grad f(\x)}^2.$$

Proof. Put $\y=\x-\alpha\grad f(\x)$ in the descent lemma: $\grad f(\x)^\top(\y-\x)=-\alpha\norm{\grad f(\x)}^2$ and $\frac L2\norm{\y-\x}^2=\frac{L\alpha^2}2\norm{\grad f(\x)}^2$. Add them. The function $\alpha(1-L\alpha/2)$ is a downward parabola in $\alpha$ with roots $0$ and $2/L$ and peak at $\alpha=1/L$, value $\frac1{2L}$. $\blacksquare$

Two connections. (1) Part 5 found that constant-step gradient descent on a quadratic is stable exactly when $0\lt\alpha\lt2/L$. The corollary extends the same threshold to every $L$-smooth function, convex or not. (2) Part 5 also showed that the gradient step of size $\alpha$ minimizes the surrogate $f(\x_k)+\g_k^\top(\y-\x_k)+\frac1{2\alpha}\norm{\y-\x_k}^2$. With $\alpha=1/L$ that surrogate is exactly the descent-lemma bound. So the $1/L$ step minimizes a function that lies above $f$ everywhere: whatever it promises, $f$ delivers at least that much. This is the "majorize, then minimize" idea.

Why this matters for line search. In practice you rarely know $L$, so you can't simply use $1/L$. But the corollary tells us that small enough steps always give a decrease proportional to $\alpha\norm{\g_k}^2$. Chapter 6.4 uses exactly this to prove that backtracking terminates with a step no smaller than a constant times $1/L$.

Try it

The green parabola is the descent-lemma bound at $x_0$ with your choice of $L$. Drag $x_0$ (or use the slider). Lower $L$ until red appears: those are points where $f$ pokes above the "bound", so that $L$ is too small. On $x^4/4$, try to find an $L$ that works on the whole window, then widen your thinking: why does no $L$ work on all of $\R$? The arrow shows the step $x_0-f'(x_0)/L$ to the parabola's lowest point, and the readout compares its guaranteed and actual decrease.

For $f(\x)=\frac12(x_1^2+9x_2^2)$ at $\x_0=(9,1)^\top$ (Tutorial 1 Q33), find $L$, take one step of size $1/L$, and compare the actual decrease with the guarantee.

  1. $\hess f=\mathrm{diag}(1,9)$, so $L=9$ (and $m=1$, $\kappa=9$).

    For a quadratic with $A\succ0$, the Lipschitz constant of the gradient is $\lambda_{\max}(A)$.

  2. $\g_0=(x_1,9x_2)^\top=(9,9)^\top$, $\norm{\g_0}^2=162$, $f(\x_0)=\frac12(81+9)=45$.

    The ingredients of the corollary.

  3. $\x_1=\x_0-\frac19\g_0=(8,0)^\top$ and $f(\x_1)=32$. Actual decrease: $13$.

    The step completely removes the steep $x_2$ component ($1-9\cdot\frac19=0$) but only shrinks $x_1$ by $\frac19$.

  4. Guarantee: $\frac1{2L}\norm{\g_0}^2=\frac{162}{18}=9\le13$. ✓

    The bound holds, and is not tight: $f$ curves less than $L$ in the $x_1$ direction, so it lies strictly below the bounding bowl there.

Find the smallest $L$ for which $f(x_1,x_2)=x_1^2+x_1x_2+2x_2^2$ is $L$-smooth.

The Hessian is constant: $\begin{pmatrix}2&1\\1&4\end{pmatrix}$. Find its largest eigenvalue from trace and determinant.

Trace $6$, determinant $7$: $\lambda^2-6\lambda+7=0$, $\lambda=3\pm\sqrt2$. Both positive (convex), so $L=\lambda_{\max}=3+\sqrt2\approx4.414$.

Find the smallest $L$ for which $f(x)=\sqrt{1+x^2}$ is $L$-smooth on $\R$.

$f'(x)=x(1+x^2)^{-1/2}$ and $f''(x)=(1+x^2)^{-3/2}$. Where is $f''$ largest?

$f''(x)=(1+x^2)^{-3/2}\in(0,1]$, largest at $x=0$. So $L=1$. (This "smoothed absolute value" is $L$-smooth even though $|x|$ is not differentiable.)

$f$ is $L$-smooth with $L=4$, and at $\x$ we have $\norm{\grad f(\x)}=3$. What decrease $f(\x)-f(\x-\alpha\grad f(\x))$ does the descent lemma guarantee for $\alpha=1/L$, and for $\alpha=0.4$?

Use $\alpha(1-L\alpha/2)\norm{\grad f}^2$.

$\alpha=0.25$: $0.25(1-0.5)\cdot9=1.125$ (equivalently $\frac{9}{2\cdot4}$). $\alpha=0.4$: $0.4(1-0.8)\cdot9=0.72$. Larger than $1/L$ but below $2/L=0.5$: still a guaranteed decrease, only a smaller one.

Which of these functions is not $L$-smooth on $\R$ for any finite $L$?

Bound $|f''|$ for each. Which second derivative grows without bound?

$|f''|$ is at most $\tfrac14$ for $\ln(1+e^x)$, $1$ for $\sin x$, $1$ for $\sqrt{1+x^2}$, but $12x^2$ for $x^4$, which is unbounded.

  • Find $L$ as the largest $|$eigenvalue$|$ of the Hessian over all points (for quadratics: $\lambda_{\max}(A)$)
  • Prove the descent lemma with: segment, integral Taylor, Cauchy–Schwarz, Lipschitz, $\int_0^1t\,dt=\frac12$
  • Remember the guaranteed decrease $\alpha(1-L\alpha/2)\norm{\g}^2$ and its peak at $\alpha=1/L$
  • Calling $x^2$ "Lipschitz" (its gradient is; the function itself is not on $\R$)
  • Assuming the descent lemma needs convexity (it doesn't)
  • Reading the lemma as a lower bound: $f$ can be far below the parabola
  • Using the asymptotic Taylor form in the proof (it only holds for small steps; the integral form is exact)
  1. $L$-smooth means $\norm{\grad f(\x)-\grad f(\y)}\le L\norm{\x-\y}$; for $C^2$ functions, curvature bounded by $L$ in absolute value.
  2. Descent lemma: $f(\y)\le f(\x)+\grad f(\x)^\top(\y-\x)+\frac L2\norm{\y-\x}^2$, proved by integral Taylor + Cauchy–Schwarz + Lipschitz.
  3. A gradient step of size $\alpha$ decreases $f$ by at least $\alpha(1-L\alpha/2)\norm{\g}^2$; at $\alpha=1/L$ that is $\frac1{2L}\norm{\g}^2$.

In the proof of the descent lemma, which step uses the $L$-Lipschitz property?

Writing $f(\y)-f(\x)$ as an integral
That's the integral form of Taylor's theorem, which needs only $f\in C^1$.
Bounding $\norm{\grad f(\z(t))-\grad f(\x)}$ by $Lt\norm{\y-\x}$
Yes: the distance from $\z(t)$ to $\x$ is $t\norm{\y-\x}$, and the gradient can change by at most $L$ times that.
Computing $\int_0^1t\,dt=\frac12$
That is plain calculus; the Lipschitz constant must enter earlier.

$f$ is $L$-smooth but not convex. Which statement is true?

The descent lemma fails, since it needs convexity
Look at the proof again: where would convexity have been used?
$f$ lies above its tangent plane everywhere
That is the convexity inequality, which a non-convex $f$ violates somewhere.
$f$ lies below the parabola $f(\x)+\grad f(\x)^\top(\y-\x)+\frac L2\norm{\y-\x}^2$
The descent lemma holds for every $L$-smooth function.

For an $L$-smooth $f$, which constant step maximizes the descent lemma's guaranteed decrease $\alpha(1-L\alpha/2)\norm{\g}^2$?

$\alpha=2/L$
At $2/L$ the factor $1-L\alpha/2$ is zero: no guarantee at all.
$\alpha=1/L$
The parabola $\alpha-\frac L2\alpha^2$ peaks where its derivative $1-L\alpha$ vanishes.
$\alpha=1/(2L)$
That gives $\frac{3}{8L}\norm{\g}^2$; can you do better?

The fit-the-line loss has Hessian eigenvalues $\approx0.721$ and $\approx33.28$. Its smallest Lipschitz constant $L$ for the gradient is about…

$0.721$
That is $m$, the smallest curvature. The gradient can change as fast as the largest curvature allows.
$46.2$
That is $\kappa=L/m$, a ratio, not a curvature.
$33.28$
$L=\lambda_{\max}(A)$ for a convex quadratic.

The acceptance tests: Armijo, Goldstein and Wolfe

Each test draws one or two straight lines through the point $(0,\phi(0))$ in the $\phi$ picture, and accepts a step $\alpha$ according to where $\phi(\alpha)$, or its slope, sits relative to those lines.

These four conditions appear in the hypotheses of the convergence theorems in Part 7 and of quasi-Newton methods in Part 10. Exams ask you to state them, draw them, check a given step, and explain why each constant must lie where it does.

A bouncer with two rules: you must have dropped enough height to come in (not too long a step), and you can't still be sliding steeply downhill when you stop (not too short).

Sufficient decrease (Armijo)

The tangent line at $0$, $\phi(0)+\alpha\phi'(0)$, is what $f$ would do if it stayed linear. Real functions curve up, so we can't ask for that much decrease. Armijo asks for a fixed fraction $c_1$ of it.

Fix $c_1\in(0,1)$ (typically $c_1=10^{-4}$). A step $\alpha\gt0$ satisfies Armijo if $$f(\x_k+\alpha\d_k)\le f(\x_k)+c_1\alpha\,\g_k^\top\d_k,\qquad\text{i.e.}\qquad \phi(\alpha)\le\phi(0)+c_1\alpha\phi'(0).$$

The right side is a line of slope $c_1\phi'(0)$: negative, but much shallower than the tangent. Acceptable steps are those where the graph of $\phi$ lies on or below this line. The decrease achieved is at least $c_1\alpha|\phi'(0)|$: proportional to the step. That proportionality is what kills the "too long" failure of Chapter 6.1: overshooting steps barely decrease $f$, so they fail.

Small steps always pass. By Taylor, $\phi(\alpha)=\phi(0)+\alpha\phi'(0)+o(\alpha)$. Since $c_1\lt1$ and $\phi'(0)\lt0$, the gap $(1-c_1)\alpha|\phi'(0)|$ between the Armijo line and the tangent beats the $o(\alpha)$ error for all small enough $\alpha$. So Armijo accepts an entire interval $(0,\bar\alpha]$. This is good news (an acceptable step always exists) and bad news: Armijo is a ceiling, never a floor. It happily accepts uselessly tiny steps.

Why $c_1\le\frac12$ is usual. If $\phi$ is a parabola $\phi(0)+\phi'(0)\alpha+\frac a2\alpha^2$ ($a\gt0$), its minimizer $\alpha^\star=-\phi'(0)/a$ gives $\phi(\alpha^\star)=\phi(0)+\frac12\alpha^\star\phi'(0)$. So the exact minimizer passes Armijo exactly when $c_1\le\frac12$. Near a minimizer most smooth $\phi$ look like parabolas, so $c_1\gt\frac12$ would reject the very steps that Newton-type methods produce.

A floor: Goldstein's second line

Fix $c\in(0,\frac12)$. A step $\alpha$ satisfies the Goldstein conditions if $$\phi(0)+(1-c)\alpha\phi'(0)\ \le\ \phi(\alpha)\ \le\ \phi(0)+c\alpha\phi'(0).$$

The right inequality is Armijo with $c_1=c$. The left one adds a steeper line of slope $(1-c)\phi'(0)$ and demands $\phi(\alpha)$ stay above it. Near $\alpha=0$, $\phi$ hugs its tangent, which is below the steep line, so tiny steps now fail: we have a floor. The requirement $c\lt\frac12$ makes the shallow line lie above the steep one ($c\lt1-c$), so there is room between them, and (by the parabola computation above) it lets the minimizer of a quadratic pass. Its drawback ([NW] p.41): on a non-quadratic $\phi$ the lower line can cut off every minimizer of $\phi$.

A floor using the slope: the Wolfe conditions

Instead of a second line, look at the slope of $\phi$ where you stop. If $\phi'(\alpha)$ is still strongly negative, $f$ is still dropping fast, so you stopped too early. If the slope has flattened (or turned positive), there is little more to gain along this ray.

Fix $0\lt c_1\lt c_2\lt1$. The curvature condition is $$\grad f(\x_k+\alpha\d_k)^\top\d_k\ \ge\ c_2\,\g_k^\top\d_k,\qquad\text{i.e.}\qquad\phi'(\alpha)\ge c_2\phi'(0).$$ The Wolfe conditions are Armijo plus curvature. The strong Wolfe conditions are Armijo plus the two-sided version $$|\phi'(\alpha)|\le c_2|\phi'(0)|.$$ Typical choices: $c_2=0.9$ for Newton and quasi-Newton directions, $c_2=0.1$ for nonlinear conjugate gradients.

Since $\phi'(0)\lt0$, the curvature condition says the slope at $\alpha$ is "less negative" than $c_2$ times the starting slope. Strong Wolfe also forbids $\phi'(\alpha)$ from being too positive, which keeps $\alpha$ near a stationary point of $\phi$ rather than far up the far side of a valley.

RuleUpper test (not too long)Lower test (not too short)Constants
Armijo$\phi(\alpha)\le\phi(0)+c_1\alpha\phi'(0)$none$0\lt c_1\lt1$
Goldstein$\phi(\alpha)\le\phi(0)+c\alpha\phi'(0)$$\phi(\alpha)\ge\phi(0)+(1-c)\alpha\phi'(0)$$0\lt c\lt\frac12$
WolfeArmijo$\phi'(\alpha)\ge c_2\phi'(0)$$0\lt c_1\lt c_2\lt1$
Strong WolfeArmijo and $\phi'(\alpha)\le c_2|\phi'(0)|$$\phi'(\alpha)\ge-c_2|\phi'(0)|$$0\lt c_1\lt c_2\lt1$

Fletcher [FR] §2.5 writes $\rho$ for $c_1$ (and for Goldstein's $c$) and $\sigma$ for $c_2$, and calls Wolfe's pair the Wolfe–Powell conditions. Nocedal–Wright [NW] uses $c_1,c_2$ as here.

Try it

Pick a slice $\phi$ and a rule. Drag on the plot (or use the slider) to move the trial step $\alpha$; the readout checks every inequality with numbers. The coloured bars under the curve mark the acceptable steps for each rule. On the lab slice with $c_1=0.05$, $c_2=0.2$, confirm: Armijo accepts tiny steps near $0$, Wolfe removes them, strong Wolfe also removes a band around $\alpha\approx1.3$ where $\phi$ climbs steeply. Then set $c_2\lt c_1$ and see what the existence theorem below is about.

Let $f\in C^1$, $\d_k$ a descent direction, and $\phi(\alpha)=f(\x_k+\alpha\d_k)$ bounded below for $\alpha\gt0$. If $0\lt c_1\lt c_2\lt1$, there are intervals of step lengths satisfying the Wolfe conditions and the strong Wolfe conditions.

Proof. (1) The Armijo line $\ell(\alpha)=\phi(0)+c_1\alpha\phi'(0)$ decreases without bound, while $\phi$ is bounded below. Near $0$, $\phi\lt\ell$ (small steps pass Armijo). So the graph of $\phi$ must cross $\ell$; let $\alpha'\gt0$ be the first crossing, $\phi(\alpha')=\phi(0)+c_1\alpha'\phi'(0)$. Every $\alpha\in(0,\alpha')$ satisfies Armijo strictly.
(2) By the mean value theorem there is $\alpha''\in(0,\alpha')$ with $\phi(\alpha')-\phi(0)=\alpha'\phi'(\alpha'')$. Comparing with (1): $\phi'(\alpha'')=c_1\phi'(0)$.
(3) Since $c_1\lt c_2$ and $\phi'(0)\lt0$: $\phi'(\alpha'')=c_1\phi'(0)\gt c_2\phi'(0)$, so curvature holds strictly at $\alpha''$. Also $\phi'(\alpha'')\lt0$, so $|\phi'(\alpha'')|=c_1|\phi'(0)|\lt c_2|\phi'(0)|$: strong Wolfe holds too. All the inequalities are strict and $\phi,\phi'$ are continuous, so they still hold on a small interval around $\alpha''$. $\blacksquare$

Notice where $c_1\lt c_2$ was used: in step (3). With $c_2\le c_1$, the slope $c_1\phi'(0)$ found by the mean value theorem would not be good enough, and on some functions no step satisfies both conditions.

Let $\phi(\alpha)=4-4\alpha+\alpha^2$ (so $\phi'(0)=-4$, minimizer $\alpha^\star=2$). With $c_1=0.1$ and $c_2=0.5$, find the steps accepted by Armijo, Goldstein ($c=0.1$), Wolfe and strong Wolfe.

  1. Armijo: $4-4\alpha+\alpha^2\le4-0.4\alpha\iff\alpha^2\le3.6\alpha\iff0\lt\alpha\le3.6$.

    For a parabola in general: $\alpha\le2(1-c_1)\alpha^\star$. It includes every tiny step.

  2. Goldstein lower line: $4-4\alpha+\alpha^2\ge4-3.6\alpha\iff\alpha\ge0.4$. With the upper line: $\alpha\in[0.4,\,3.6]$.

    In general $[2c\,\alpha^\star,\ 2(1-c)\alpha^\star]$, symmetric around $\alpha^\star$. The floor $0.4$ excludes tiny steps.

  3. Curvature: $\phi'(\alpha)=-4+2\alpha\ge0.5(-4)=-2\iff\alpha\ge1$. Wolfe: $\alpha\in[1,\,3.6]$.

    In general $\alpha\ge(1-c_2)\alpha^\star$: the slope has flattened to half its starting value.

  4. Strong Wolfe: $|-4+2\alpha|\le2\iff1\le\alpha\le3$; with Armijo, $\alpha\in[1,\,3]$.

    The two-sided slope test also cuts the far side where $\phi$ is climbing fast: $[(1-c_2)\alpha^\star,(1+c_2)\alpha^\star]$.

  5. All four sets contain the exact minimizer $\alpha^\star=2$.

    Here $c_1=c=0.1\le\frac12$, as the parabola argument requires.

Back to the trap

For $f(x)=x^2$ with $\d_k=-2x_k$: $\phi(\alpha)=x_k^2(1-2\alpha)^2$ and $\phi'(\alpha)=-4x_k^2(1-2\alpha)$. Armijo reduces to $\alpha\le1-c_1$, so every short step $\alpha_k=2^{-(k+2)}$ passes, and the long steps $1-2^{-(k+2)}$ pass until $2^{-(k+2)}\lt c_1$. Curvature reduces to $1-2\alpha\le c_2$, i.e. $\alpha\ge(1-c_2)/2$. With $c_2=0.9$ that is $\alpha\ge0.05$: $\alpha_0,\alpha_1,\alpha_2$ pass but $\alpha_3=\frac1{32}=0.03125$ fails, and so does every later short step. The Wolfe line search would refuse them and look for a larger step. Strong Wolfe gives $0.05\le\alpha\le0.95$ and also rejects the long steps from $\alpha_3=\frac{31}{32}$ on.

$\phi(\alpha)=12-6\alpha+\alpha^2$. Find the largest step satisfying Armijo with $c_1=0.1$, and the interval of steps satisfying Goldstein with $c=0.25$.

$\phi'(0)=-6$, $\alpha^\star=3$. Armijo: $\alpha^2-6\alpha\le-0.6\alpha$. Goldstein lower line has slope $(1-c)\phi'(0)=-4.5$.

Armijo: $\alpha^2\le5.4\alpha\iff\alpha\le5.4$ ($=2(1-c_1)\alpha^\star$). Goldstein: upper line $12-1.5\alpha$ gives $\alpha\le4.5$; lower line $12-4.5\alpha$ gives $\alpha^2-6\alpha\ge-4.5\alpha\iff\alpha\ge1.5$. Interval $[1.5,\,4.5]=[2c\alpha^\star,2(1-c)\alpha^\star]$.

Same $\phi(\alpha)=12-6\alpha+\alpha^2$, $c_1=0.1$. Find the smallest step satisfying the curvature condition with $c_2=0.9$, and the strong Wolfe interval with $c_2=0.5$.

$\phi'(\alpha)=-6+2\alpha$. Curvature: $\phi'(\alpha)\ge c_2(-6)$. Strong: $|{-6}+2\alpha|\le6c_2$. Don't forget to intersect with Armijo ($\alpha\le5.4$).

$c_2=0.9$: $-6+2\alpha\ge-5.4\iff\alpha\ge0.3$. Strong with $c_2=0.5$: $|2\alpha-6|\le3\iff1.5\le\alpha\le4.5$, which lies inside the Armijo range, so the interval is $[1.5,\,4.5]$.

The lab slice: $f(x)=x^4-3x^2+x+4$, $x_k=-2$, $d=1$, so $\phi(\alpha)=f(-2+\alpha)$ with $\phi(0)=6$, $\phi'(0)=-19$. With $c_1=0.05$, $c_2=0.2$, test the step $\alpha=1.3$.

The point is $x=-0.7$. Compute $f(-0.7)$ and $f'(-0.7)=4x^3-6x+1$, then compare with $6-0.05\cdot19\cdot1.3$ and with $\pm0.2\cdot19=\pm3.8$.

$\phi(1.3)=f(-0.7)=0.2401-1.47-0.7+4=2.0701\le6-1.235=4.765$: Armijo holds. $\phi'(1.3)=f'(-0.7)=-1.372+4.2+1=3.828\ge-3.8$: curvature holds, so Wolfe holds. But $|3.828|\gt3.8$: strong Wolfe fails, because $\phi$ is climbing steeply there.

For every strictly convex quadratic $\phi(\alpha)=\phi(0)+\phi'(0)\alpha+\frac a2\alpha^2$ ($a\gt0$, $\phi'(0)\lt0$), what is the largest $c_1$ for which the exact minimizer always satisfies Armijo?

Compute $\phi(\alpha^\star)$ with $\alpha^\star=-\phi'(0)/a$, and write it as $\phi(0)+(\text{something})\cdot\alpha^\star\phi'(0)$.

$\phi(\alpha^\star)=\phi(0)-\frac{\phi'(0)^2}{2a}=\phi(0)+\frac12\alpha^\star\phi'(0)$. Armijo asks $\frac12\alpha^\star\phi'(0)\le c_1\alpha^\star\phi'(0)$; dividing by the negative number $\alpha^\star\phi'(0)$ flips it to $c_1\le\frac12$. So $c_1=0.5$.

  • Draw the picture: $\phi$, its tangent, the Armijo line of slope $c_1\phi'(0)$, and (Goldstein) the line of slope $(1-c)\phi'(0)$
  • Write curvature as $\phi'(\alpha)\ge c_2\phi'(0)$ and remember $\phi'(0)\lt0$ when manipulating it
  • State the constants: $0\lt c_1\lt c_2\lt1$ for Wolfe, $0\lt c\lt\frac12$ for Goldstein
  • For quadratic $\phi$, use the shortcuts $2(1-c_1)\alpha^\star$, $(1\pm c_2)\alpha^\star$, $[2c,2(1-c)]\alpha^\star$
  • Writing curvature as $|\phi'(\alpha)|\ge c_2|\phi'(0)|$ (the inequality goes the other way, and only strong Wolfe uses absolute values)
  • Flipping an inequality wrongly when dividing by $\phi'(0)\lt0$
  • Thinking Armijo alone stops steps from being too small
  • Thinking acceptable steps must be near the nearest minimizer of $\phi$ (in the lab slice they lie in both wells)
  1. Armijo $\phi(\alpha)\le\phi(0)+c_1\alpha\phi'(0)$ demands decrease proportional to the step: it caps steps from above and always accepts small ones.
  2. Goldstein (second line, $c\lt\frac12$) and Wolfe (slope test $\phi'(\alpha)\ge c_2\phi'(0)$) add a floor; strong Wolfe also caps the slope from above.
  3. With $0\lt c_1\lt c_2\lt1$ and $\phi$ bounded below, (strong) Wolfe steps always exist: first Armijo crossing + mean value theorem.

Why must $c_1\lt1$ in the Armijo condition?

So that the Armijo line is steeper than the tangent
Slope $c_1\phi'(0)$ with $c_1\lt1$ is shallower than the tangent's slope $\phi'(0)$. Which of the two lines is $\phi$ below for small $\alpha$?
So that small steps satisfy it: near $0$, $\phi$ follows its tangent, which lies below the shallower Armijo line
With $c_1\ge1$, a function curving upward ($\phi''(0)\gt0$) would fail Armijo for all small $\alpha\gt0$.
So that the exact minimizer of a quadratic is accepted
That needs the stronger restriction $c_1\le\frac12$.

At a trial step, $\phi'(0)=-10$ and $\phi'(\alpha)=-9.5$. With $c_2=0.9$, the curvature condition…

holds, because $-9.5\gt-10$
The comparison is with $c_2\phi'(0)$, not with $\phi'(0)$.
fails, because $-9.5\lt c_2\phi'(0)=-9$
$\phi$ is still descending almost as steeply as at the start, so the step stopped too early.
cannot be checked without $\phi(\alpha)$
The curvature condition only involves slopes.

In the proof that Wolfe steps exist, where is $c_1\lt c_2$ used?

To show that the Armijo line crosses the graph of $\phi$
That uses $\phi$ bounded below and $c_1\gt0$.
To apply the mean value theorem
The MVT needs only differentiability.
To conclude that $\phi'(\alpha'')=c_1\phi'(0)$ exceeds $c_2\phi'(0)$
Multiplying $c_1\lt c_2$ by $\phi'(0)\lt0$ reverses it: $c_1\phi'(0)\gt c_2\phi'(0)$.

Which statement about Goldstein vs Wolfe is correct?

Goldstein needs gradient evaluations at trial points
Both Goldstein lines involve only $\phi(\alpha)$ and the starting slope.
Goldstein's lower line can exclude every minimizer of a non-quadratic $\phi$; the Wolfe slope test cannot exclude a local minimizer that satisfies Armijo
At a local minimizer of $\phi$, $\phi'=0\ge c_2\phi'(0)$, so curvature holds automatically.
Goldstein allows $c$ up to $1$
With $c\ge\frac12$ the two lines cross over and the sandwich can be empty.

Backtracking line search, and why it never stalls

Start with a bold step, and halve it (or shrink it by a fixed factor) until Armijo is satisfied. That's the whole algorithm.

It is the line search inside most gradient-descent codes: it needs only function values, no gradients at trial points. And although it checks only Armijo, it provably avoids the too-short trap. That proof is an exam favourite.

Throwing a ball to someone and walking closer each time it falls short of their hands: you start far away, and you only step in until it works, never further.

Input: $\x_k$, a descent direction $\d_k$, initial step $\bar\alpha\gt0$ (often $1$), contraction factor $\beta\in(0,1)$ (often $0.5$), $c_1\in(0,1)$ (often $10^{-4}$).

  1. Set $\alpha\leftarrow\bar\alpha$.
  2. While $f(\x_k+\alpha\d_k)\gt f(\x_k)+c_1\alpha\,\g_k^\top\d_k$: set $\alpha\leftarrow\beta\alpha$.
  3. Return $\alpha_k=\alpha$.

[NW] calls the contraction factor $\rho$. The trials are $\bar\alpha,\beta\bar\alpha,\beta^2\bar\alpha,\dots$, and the first one that passes Armijo is accepted.

(Tutorial 1 Q30.) $f(x)=x^4-3x^2+2$, $x_0=1.5$, $d_0=-f'(x_0)$, $\bar\alpha=1$, $\beta=0.5$, $c_1=0.1$. Run backtracking.

  1. $f(1.5)=5.0625-6.75+2=0.3125$; $f'(x)=4x^3-6x$, so $f'(1.5)=13.5-9=4.5$ and $d_0=-4.5$. Slope $\phi'(0)=f'(x_0)d_0=-20.25$.

    Everything Armijo needs is computed once, before the loop.

  2. Armijo threshold: $0.3125+0.1\alpha(-20.25)=0.3125-2.025\alpha$.

    A straight line in $\alpha$; each trial compares $f(x_0+\alpha d_0)$ with it.

  3. $\alpha$$x_0+\alpha d_0$$f$thresholdaccept?
    1−356−1.7125no
    0.5−0.750.6289−0.7no
    0.250.3751.5979−0.1938no
    0.1250.93750.13580.0594no
    0.06251.21875−0.24980.1859yes

    Each rejection halves $\alpha$. Note trial 4: $f$ did go down ($0.1358\lt0.3125$), but not by enough for a step of that length.

  4. $\alpha_0=0.0625$, $x_1=1.21875$, $f(x_1)\approx-0.2498$, close to the minimum value $-0.25$ at $x=\sqrt{1.5}\approx1.2247$.

    Five function evaluations, no derivative at any trial point, and a near-optimal step.

Why it terminates

Only two facts are needed: small steps pass Armijo (Chapter 6.3, using $c_1\lt1$ and $\phi'(0)\lt0$), and the trials $\beta^j\bar\alpha$ go to $0$. So some trial is small enough, and the loop stops after finitely many trials. This works for any $f\in C^1$.

But termination alone does not stop the "too short" failure: what if the accepted step is absurdly small? For $L$-smooth $f$ the descent lemma rules that out, with an explicit bound. The idea: the step just before the accepted one was rejected, and a rejection tells you that $f$ curves up a lot, which the descent lemma says can only happen for steps of size at least about $1/L$.

Let $f$ be $L$-smooth and $\d_k$ a descent direction. Backtracking terminates after finitely many trials with $$\alpha_k\ \ge\ \min\left\{\bar\alpha,\ \frac{2\beta(1-c_1)\,|\g_k^\top\d_k|}{L\norm{\d_k}^2}\right\}.$$ For steepest descent ($\d_k=-\g_k$) this is $\alpha_k\ge\min\{\bar\alpha,\ 2\beta(1-c_1)/L\}$, and the accepted step decreases $f$ by at least $$f(\x_k)-f(\x_{k+1})\ \ge\ c_1\min\Big\{\bar\alpha,\ \frac{2\beta(1-c_1)}L\Big\}\norm{\g_k}^2.$$

Prove the theorem. (Examinable: Book Ch. 11 Examtip; Tutorial 1 Q32.)

  1. Case 1: $\bar\alpha$ is accepted at once. Then $\alpha_k=\bar\alpha$, which is at least the minimum.

    Always dispose of the trivial case first.

  2. Case 2: at least one rejection. Then $t:=\alpha_k/\beta$ was the last rejected trial, so Armijo failed at $t$: $$f(\x_k+t\d_k)\gt f(\x_k)+c_1t\,\g_k^\top\d_k=f(\x_k)-c_1tG,\qquad G:=-\g_k^\top\d_k\gt0.$$

    Key insight: a rejection is information. It gives a lower bound on $f$ at the point $\x_k+t\d_k$.

  3. The descent lemma with $\y=\x_k+t\d_k$ gives an upper bound at the same point: $$f(\x_k+t\d_k)\le f(\x_k)-tG+\tfrac L2t^2\norm{\d_k}^2.$$

    Valid for every $t\ge0$ because $f$ is $L$-smooth.

  4. Chain them: $f(\x_k)-c_1tG\lt f(\x_k)-tG+\tfrac L2t^2\norm{\d_k}^2$, i.e. $(1-c_1)tG\lt\tfrac L2t^2\norm{\d_k}^2$. Divide by $t\gt0$: $$t\gt\frac{2(1-c_1)G}{L\norm{\d_k}^2}.$$

    The lower bound from the rejection can sit below the upper bound from the lemma only if $t$ is large enough.

  5. $\alpha_k=\beta t\gt\dfrac{2\beta(1-c_1)G}{L\norm{\d_k}^2}$. Combining the cases gives the bound. For $\d_k=-\g_k$, $G=\norm{\g_k}^2=\norm{\d_k}^2$, so the fraction is $2\beta(1-c_1)/L$; inserting $\alpha_k$ into Armijo, $f(\x_k)-f(\x_{k+1})\ge c_1\alpha_k\norm{\g_k}^2$, gives the decrease bound. $\blacksquare$

    The bound depends only on $L,\beta,c_1,\bar\alpha$ and the direction, never on $k$: the steps cannot shrink towards $0$ as iterations go on.

From the lower bound to convergence

Write $\omega=c_1\min\{\bar\alpha,2\beta(1-c_1)/L\}\gt0$. If $f$ is bounded below by $f^\star$, add the decrease bound over $k=0,\dots,N-1$; the left side telescopes: $$\omega\sum_{k=0}^{N-1}\norm{\g_k}^2\le f(\x_0)-f(\x_N)\le f(\x_0)-f^\star.$$ The right side doesn't depend on $N$, so the series $\sum\norm{\g_k}^2$ converges and $\norm{\g_k}\to0$. Backtracking gradient descent is globally convergent to stationary points, without any curvature condition. This resolves the summable-step trap: an externally imposed schedule $2^{-(k+2)}$ can shrink forever, but backtracking restarts from $\bar\alpha$ every iteration and only shrinks while Armijo actually fails. Part 7 extends this kind of argument to Wolfe steps and general directions (Zoutendijk).

A practical by-product: any $\alpha\le2(1-c_1)G/(L\norm{\d_k}^2)$ passes Armijo (that is what step 4 shows in reverse), so the number of trials is at most the number of halvings needed to get below that value, plus one.

Try it

Press "Next trial" to watch backtracking shrink $\alpha$ until the trial point (on the map) and the trial dot (on $\phi$) fall below the Armijo line; the $\phi$ window zooms in on the current trial. Then "Next iteration" to start a new search from the new point, or "Run 20 iterations". On the stretched bowl, compare the accepted steps with the guaranteed floor $\min\{\bar\alpha,2\beta(1-c_1)/L\}$. Try $\beta=0.9$ (many gentle shrinks) against $\beta=0.1$ (few, coarse ones), and a large $c_1$ such as $0.8$.

$f(x)=e^x-x-1$, $x_0=1$, $d_0=-f'(x_0)$, $c_1=0.2$, $\bar\alpha=1$, $\beta=0.5$. Find the accepted backtracking step.

$f(1)=e-2\approx0.71828$, $f'(1)=e-1\approx1.71828$, $\phi'(0)=-(e-1)^2\approx-2.9525$. Threshold: $0.71828-0.5905\alpha$. Try $\alpha=1$ (point $x\approx-0.71828$) first.

$\alpha=1$: $f(-0.71828)\approx0.48759+0.71828-1=0.20587$, threshold $0.12778$: reject. $\alpha=0.5$: point $0.14086$, $f\approx1.15126-1.14086=0.01040$, threshold $0.42303$: accept. So $\alpha_0=0.5$.

$f(\x)=\frac12(x_1^2+9x_2^2)$, $\x_0=(9,1)^\top$, $\d_0=-\g_0$, $\bar\alpha=1$, $\beta=0.5$, $c_1=10^{-4}$. Find the accepted step and $f(\x_1)$.

$\g_0=(9,9)^\top$, $f(\x_0)=45$, $\norm{\g_0}^2=162$. At $\alpha=1$ the point is $(0,-8)$.

$\alpha=1$: $(0,-8)$, $f=288\gt45-0.0162$: reject. $\alpha=0.5$: $(4.5,-3.5)$, $f=\frac12(20.25+110.25)=65.25$: reject. $\alpha=0.25$: $(6.75,-1.25)$, $f=\frac12(45.5625+14.0625)=29.8125\le45-0.00405$: accept. Compare the floor $\min\{1,2(0.5)(0.9999)/9\}\approx0.111$: indeed $0.25\ge0.111$.

Backtracking steepest descent on an $L$-smooth $f$ with $L=10$, $\bar\alpha=1$, $\beta=0.5$, $c_1=0.1$. At an iterate with $\norm{\g_k}=2$, what lower bound does the theorem give on $\alpha_k$, and on the decrease $f(\x_k)-f(\x_{k+1})$?

$\alpha_k\ge\min\{\bar\alpha,2\beta(1-c_1)/L\}$; decrease $\ge c_1\alpha_{\min}\norm{\g_k}^2$.

$2(0.5)(0.9)/10=0.09\lt1$, so $\alpha_k\ge0.09$. Decrease $\ge0.1\cdot0.09\cdot4=0.036$.

Steepest descent on an $L$-smooth $f$ with $L=100$, backtracking with $\bar\alpha=1$, $\beta=0.5$, $c_1=10^{-4}$. What is the largest possible number of trial steps (function evaluations at trial points) in one line search?

Every $\alpha\le2(1-c_1)/L\approx0.019998$ passes Armijo. Which is the first trial $2^{-j}$ at or below it?

$2^{-5}=0.03125\gt0.019998$ but $2^{-6}=0.015625\le0.019998$, so trial $\alpha=2^{-6}$ is certainly accepted (if no earlier one was). The trials are $2^{0},\dots,2^{-6}$: at most $7$.

  • Compute $f(\x_k)$ and $\g_k^\top\d_k$ once, then tabulate trials: $\alpha$, point, $f$, threshold, verdict
  • Prove the lower bound by chaining "rejected at $t=\alpha_k/\beta$" with the descent lemma at the same point
  • Restart from $\bar\alpha$ at every iteration (that is what makes the floor independent of $k$)
  • Accepting a trial because $f$ decreased, without comparing to the Armijo threshold
  • Forgetting the factor $\beta$ in the bound (the accepted step is $\beta$ times the last rejected one)
  • Using $\alpha_k$ itself in the rejection inequality (it was accepted; the rejected one is $\alpha_k/\beta$)
  • Carrying the previous $\alpha_k$ over as the next $\bar\alpha$ without thinking: steps could then only shrink
  1. Backtracking tries $\bar\alpha,\beta\bar\alpha,\beta^2\bar\alpha,\dots$ and accepts the first step passing Armijo; it terminates for any $C^1$ $f$ and descent direction.
  2. For $L$-smooth $f$: $\alpha_k\ge\min\{\bar\alpha,2\beta(1-c_1)G/(L\norm{\d_k}^2)\}$, proved by chaining the last rejection with the descent lemma.
  3. For steepest descent this gives a decrease $\ge c_1\min\{\bar\alpha,2\beta(1-c_1)/L\}\norm{\g_k}^2$, so $\norm{\g_k}\to0$ when $f$ is bounded below.

Backtracking checks only Armijo, which has no floor. Why can't its accepted steps shrink to zero over the iterations (for $L$-smooth $f$)?

Because $\beta\lt1$ makes the steps shrink geometrically
Shrinking geometrically is the opposite of having a floor. What does a rejected trial tell you?
Because the last rejected step $\alpha_k/\beta$ must exceed $2(1-c_1)G/(L\norm{\d_k}^2)$, so $\alpha_k$ is at least $\beta$ times that
Steps below the descent-lemma threshold are never rejected, so the search stops before going much below it.
Because Armijo implies the curvature condition
It doesn't: the trap schedule satisfies Armijo and fails curvature.

Backtracking with $\bar\alpha=1$, $\beta=0.5$ accepts $\alpha_k=0.125$. Which inequality does the proof use?

$f(\x_k+0.125\d_k)\gt f(\x_k)+c_1(0.125)\g_k^\top\d_k$
$0.125$ was accepted, so Armijo holds there.
$f(\x_k+\d_k)\le f(\x_k)+c_1\g_k^\top\d_k$
$\alpha=1$ was rejected, so this inequality is false.
$f(\x_k+0.25\d_k)\gt f(\x_k)+c_1(0.25)\g_k^\top\d_k$
The last rejected trial is $\alpha_k/\beta=0.25$; its failed Armijo test is chained with the descent lemma.

Each trial of backtracking costs…

one function evaluation; the gradient is needed only once, at $\x_k$
The threshold uses $\g_k^\top\d_k$, computed before the loop. Wolfe searches also need $\grad f$ at trial points.
one function and one gradient evaluation
Does the Armijo test involve the gradient at the trial point?
a Hessian evaluation
No second derivatives appear anywhere in Armijo.

Steepest descent with backtracking, $L=4$, $\bar\alpha=1$, $\beta=0.5$, $c_1=0.5$. The guaranteed floor on $\alpha_k$ is…

$1/L=0.25$
The floor is $\min\{\bar\alpha,2\beta(1-c_1)/L\}$; put in the numbers.
$0.125$
$2(0.5)(0.5)/4=0.125\lt1$.
$0.5$
That would be $2(1-c_1)/L$ without the factor $\beta$.

Which form of Taylor's theorem does the descent lemma's proof use, and why?

Asymptotic form, because $\y$ is close to $\x$
The lemma holds for all $\x,\y$, however far apart, so a small-step approximation can't prove it.
Integral form, because it is exact for any distance and needs only $f\in C^1$
Then Cauchy–Schwarz and the Lipschitz bound act inside the integral.
Lagrange form, because it involves the Hessian
$L$-smooth functions need not have a Hessian at all; the proof uses only gradients.

Part 5 showed constant-step gradient descent on a quadratic is stable iff $0\lt\alpha\lt2/L$. For a general $L$-smooth $f$, the descent lemma guarantees decrease for…

only $\alpha=1/L$
$1/L$ gives the best guarantee, but look at the sign of $\alpha(1-L\alpha/2)$ for other $\alpha$.
every $\alpha\gt0$, since $\d_k$ is a descent direction
Descent direction only guarantees decrease for small enough steps.
every $\alpha\in(0,2/L)$, matching the quadratic case
$\alpha(1-L\alpha/2)\gt0$ exactly on $(0,2/L)$.

At $\x_k$, $\g_k=(2,-2)^\top$. For which direction does every line-search rule of this part make sense?

$\d=(1,1)^\top$
$\g_k^\top\d=0$: flat, not downhill, so $\phi'(0)\lt0$ fails.
$\d=(-1,0)^\top$
$\g_k^\top\d=-2\lt0$: a descent direction.
$\d=(1,-1)^\top$
$\g_k^\top\d=4\gt0$: uphill.

On $\phi(\alpha)=9-6\alpha+\alpha^2$ ($\alpha^\star=3$), with $c_1=10^{-4}$, $c_2=0.9$, which step satisfies Armijo but not Wolfe?

$\alpha=3$
At the minimizer $\phi'=0\ge c_2\phi'(0)$: curvature holds.
$\alpha=0.1$
Curvature needs $\alpha\ge(1-c_2)\alpha^\star=0.3$, while Armijo accepts everything up to $\approx6$.
$\alpha=6.5$
Beyond $2(1-c_1)\alpha^\star\approx5.9994$, Armijo fails too.

Why is backtracking gradient descent globally convergent while the schedule $\alpha_k=2^{-(k+2)}$ (which satisfies Armijo every time) is not?

Backtracking uses a smaller $c_1$
The schedule passes Armijo for any $c_1\lt0.75$; $c_1$ is not the issue.
Backtracking computes the exact step
It stops at the first Armijo step, usually far from exact.
Backtracking's accepted steps stay above a fixed floor, so the decreases $\ge\omega\norm{\g_k}^2$ force $\norm{\g_k}\to0$
The imposed schedule is summable and has no floor; backtracking restarts at $\bar\alpha$ each iteration.

Lectures 8–10 (1, 3 and 8 September): what it means for an algorithm to converge and how to measure its speed; Zoutendijk's Global Convergence Theorem; what a constant step size guarantees on any smooth function; the $O(1/k)$ rate for convex functions; and the linear rate for strongly convex functions. This is the last block of the midsem syllabus, and the proofs here are exam favourites.

You need: Part 4 ($L$-smoothness, convexity, strong convexity, the classes $\mathcal F_L^{1,1}$ and $\mathcal S_{\mu,L}^{1,1}$), Part 5 (gradient descent on quadratics, condition number $\kappa$), Part 6 (the Descent Lemma, Armijo, Wolfe and Goldstein conditions).

Measuring speed: sublinear, linear, superlinear, quadratic

An algorithm "converges" when its error shrinks to zero; its "rate" says how fast, and the rate decides how many iterations you pay for each extra digit of accuracy.

Every theorem in this part ends in a rate. To read them, compare methods and answer "how many iterations?" questions, you need this vocabulary first.

Paying off a loan: a fixed amount each month (sublinear-like: the last bit takes as long as the first), a fixed percentage each month (linear: steady progress on a log scale), or a payment that squares your shrinking balance (quadratic: done almost instantly at the end).

An iterative method produces $\x_0,\x_1,\x_2,\dots$. To talk about speed we pick an error measure $e_k\ge0$ that is zero exactly at the answer. Three are used in this part:

  • the optimality gap $f(\x_k)-f^\star$, where $f^\star=\inf f$;
  • the distance $\norm{\x_k-\x^\star}$ to a minimizer $\x^\star$;
  • the gradient norm $\norm{\grad f(\x_k)}$ (zero exactly at stationary points).

"The method converges" means $e_k\to0$. The rate describes how $e_k$ shrinks.

Let $e_k\to0$ with $e_k>0$.

  • Linear (also called geometric) with rate $\rho\in(0,1)$: $e_{k+1}\le\rho\,e_k$ for all $k$ (large enough). Then $e_k\le\rho^k e_0$.
  • Sublinear: slower than every linear rate, i.e. $e_{k+1}/e_k\to1$. Typical forms: $e_k\le C/k$ or $e_k\le C/\sqrt k$.
  • Superlinear: $e_{k+1}/e_k\to0$ (faster than every linear rate).
  • Quadratic: $e_{k+1}\le M e_k^2$ for some constant $M$ (a special, very fast case of superlinear).

Why "linear"? Take logarithms of $e_k=\rho^ke_0$: $\log e_k=\log e_0+k\log\rho$, a straight line in $k$. On a plot with a logarithmic error axis (a semi-log plot), linear convergence is a straight line going down; sublinear convergence is a curve that keeps flattening; quadratic convergence is a curve that bends down ever more steeply (the number of correct digits roughly doubles each step).

How many iterations for accuracy $\varepsilon$?

This is the question exams (and practitioners) actually ask. Solve $e_k\le\varepsilon$ for $k$:

RateError after $k$ stepsIterations for $e_k\le\varepsilon$Cost of one more digit
Sublinear$C/k$$k\ge C/\varepsilon$10 times the total work
Linear$\rho^ke_0$$k\ge\dfrac{\ln(e_0/\varepsilon)}{\ln(1/\rho)}$a fixed $\ln10/\ln(1/\rho)$ extra steps
Quadratic$e_k\approx(Me_0)^{2^k}/M$about $\log_2\log(1/\varepsilon)$at most one extra step

For gradient descent the linear rate is usually $\rho=\frac{\kappa-1}{\kappa+1}$ (Part 5, and Chapter 7.4 below). Since $\ln\frac1\rho=\ln\frac{\kappa+1}{\kappa-1}\approx\frac2\kappa$ for large $\kappa$, the iteration count is roughly $$k\approx\frac\kappa2\ln\frac{e_0}{\varepsilon}.$$ So the condition number multiplies the work: $\kappa=1000$ means about 500 steps per factor $e$ of improvement.

Method S has error $1/(k+1)$ (sublinear); method G has error $0.99^k$ (linear). Both start at error 1. Which one first reaches error $10^{-2}$?

G, because linear always beats sublinear
Linear wins eventually. But $0.99^k$ shrinks by only 1% per step. Count the steps.
S, after 99 steps; G needs 459
$1/(k+1)\le0.01$ at $k=99$. For G, $k\ge\ln100/\ln(1/0.99)=458.2$, so 459. At accuracy $10^{-6}$ the order flips: G needs 1375 steps, S needs about a million.
They arrive at the same time
Compute both: $1/(k+1)=0.01$ and $0.99^k=0.01$ have quite different solutions.
Try it

The error axis is logarithmic. Set $\kappa$ near 200 (so $\rho\approx0.99$) and move $\varepsilon$ from $10^{-2}$ down to $10^{-10}$: watch which curve crosses the dashed line first, and how the iteration counts grow. Switch the $k$ axis to log scale: the sublinear curve becomes a straight line of slope $-1$.

Three error sequences start at $e_0$ and must reach $\varepsilon=10^{-6}$: (a) $e_k=1/k$; (b) $e_k=0.5^k$; (c) $e_{k+1}=e_k^2$ with $e_0=0.5$. How many iterations does each need, and what kind of rate is each?

  1. (a) $1/k\le10^{-6}$ iff $k\ge10^6$: one million iterations. The ratio $e_{k+1}/e_k=k/(k+1)\to1$: sublinear.

    A ratio tending to 1 means each step removes a vanishing fraction of the error.

  2. (b) $0.5^k\le10^{-6}$ iff $k\ge\ln10^6/\ln2=19.93$, so $k=20$ ($0.5^{19}\approx1.9\times10^{-6}$, $0.5^{20}\approx9.5\times10^{-7}$). Linear with $\rho=0.5$.

    Always round up: $k$ must be an integer and the inequality must hold.

  3. (c) By induction $e_k=0.5^{2^k}$. $k=4$ gives $0.5^{16}\approx1.5\times10^{-5}$ (not enough); $k=5$ gives $0.5^{32}\approx2.3\times10^{-10}$. So 5 iterations. Quadratic with $M=1$.

    The exponent doubles each step: that is "the number of correct digits doubles".

A method satisfies $e_k=0.9^k$ with $e_0=1$. What is the smallest $k$ with $e_k\le10^{-3}$?

Solve $0.9^k\le10^{-3}$ by taking logarithms: $k\ge\ln(1000)/\ln(1/0.9)$, then round up.

$\ln1000/\ln(1/0.9)=6.9078/0.10536=65.56$. So $k=66$. Check: $0.9^{65}\approx1.06\times10^{-3}$, $0.9^{66}\approx0.955\times10^{-3}$.

Classify $e_k=1/k^2$.

Compute $e_{k+1}/e_k$ and its limit.

$e_{k+1}/e_k=k^2/(k+1)^2\to1$. A ratio tending to 1 is sublinear, however fast $1/k^2$ may look compared with $1/k$. No $\rho\lt1$ satisfies $e_{k+1}\le\rho e_k$ for all large $k$.

Classify $e_k=1/k!$.

Look at $e_{k+1}/e_k$ (does it go to 0?) and at $e_{k+1}/e_k^2$ (does it stay bounded?).

$e_{k+1}/e_k=1/(k+1)\to0$: superlinear. But $e_{k+1}/e_k^2=k!/(k+1)\to\infty$, so no constant $M$ gives $e_{k+1}\le Me_k^2$: not quadratic.

Gradient descent on a quadratic with $\kappa=100$ and step $2/(m+L)$ shrinks $\norm{\x_k-\x^\star}$ by exactly $\rho=\frac{\kappa-1}{\kappa+1}$ per step (worst case). How many steps guarantee the distance shrinks by a factor $10^6$?

$\rho=99/101$. Solve $\rho^k\le10^{-6}$. Compare with $\frac\kappa2\ln10^6$.

$k\ge\ln10^6/\ln(101/99)=13.8155/0.020001=690.75$, so $k=691$. The approximation $\frac\kappa2\ln10^6=690.8$ is almost exact.

  • Say which error you measure: $f(\x_k)-f^\star$, $\norm{\x_k-\x^\star}$ or $\norm{\grad f(\x_k)}$
  • Classify a rate by the limit of $e_{k+1}/e_k$
  • Use log-scale plots: linear rates are straight lines on semi-log axes, $C/k^p$ rates are straight lines on log–log axes
  • Round iteration counts up
  • Calling $1/k^2$ "linear" or "quadratic" because it looks fast
  • Assuming a linear method always beats a sublinear one at the accuracy you care about
  • Forgetting that a linear rate close to 1 (large $\kappa$) can be painfully slow
  1. Linear: $e_k\le\rho^ke_0$, a straight line on a semi-log plot; iterations $\approx\ln(e_0/\varepsilon)/\ln(1/\rho)\approx\frac\kappa2\ln(e_0/\varepsilon)$ when $\rho=\frac{\kappa-1}{\kappa+1}$.
  2. Sublinear ($C/k$): iterations $\approx C/\varepsilon$, so every extra digit costs ten times the work.
  3. Superlinear: $e_{k+1}/e_k\to0$; quadratic: $e_{k+1}\le Me_k^2$, correct digits double each step.

On a plot with logarithmic error axis and ordinary $k$ axis, a linearly convergent method appears as…

a curve that flattens out
That's the signature of a ratio $e_{k+1}/e_k$ creeping towards 1.
a straight line going down
$\log e_k=\log e_0+k\log\rho$ is linear in $k$ with slope $\log\rho\lt0$.
a curve bending down ever more steeply
That's superlinear behaviour: the per-step factor itself keeps shrinking.

A method has $f(\x_k)-f^\star\le C/k$. To guarantee 1000 times smaller error, you need about…

a fixed number of extra iterations
That is how linear rates behave. Solve $C/k\le\varepsilon/1000$.
about 10 times more iterations
The bound is inversely proportional to $k$. How does $k$ scale with $1/\varepsilon$?
1000 times more iterations
$k\ge C/\varepsilon$, so dividing $\varepsilon$ by 1000 multiplies $k$ by 1000.

Two linearly convergent methods have $\rho_1=0.5$ and $\rho_2=0.9$. Roughly how many steps of method 2 match one step of method 1?

2
Compare the logarithms, not the rates: how many factors of 0.9 make 0.5?
about 6.6
$0.9^n=0.5$ gives $n=\ln2/\ln(1/0.9)\approx6.58$.
about 1.8
That's $0.9/0.5$, but rates multiply; you need a ratio of logarithms.

$e_k=2^{-2^k}$ converges…

linearly with $\rho=1/2$
Check $e_{k+1}/e_k=2^{-2^k}$: does it stay at $1/2$?
sublinearly
The ratio $e_{k+1}/e_k$ goes to 0, not to 1.
quadratically
$e_{k+1}=2^{-2^{k+1}}=(2^{-2^k})^2=e_k^2$, so $M=1$.

Global convergence: Zoutendijk's theorem and constant step sizes

For almost any sensible line-search method, Zoutendijk's theorem guarantees that the gradient goes to zero from any starting point, provided the search directions never become nearly perpendicular to the gradient.

This is the theorem that lets you call steepest descent, Newton-like and quasi-Newton methods "globally convergent", and its statement (three hypotheses) and its three-step proof are standard exam material.

Walking downhill in fog: as long as each step points at least somewhat downhill and you never end up walking almost along the contour, you can't wander forever on a slope; you must end up somewhere flat.

What "global" does and does not mean

A method is globally convergent if $\norm{\grad f(\x_k)}\to0$ from every starting point $\x_0$. "Global" refers to the starting point, not to the answer: the limit is a stationary point, which may be a saddle or (rarely) a maximum. Without convexity nothing better can be promised.

Nesterov's warning example ([Y] Example 1.2.2)

Take $f(x,y)=\tfrac12x^2+\tfrac14y^4-\tfrac12y^2$. Its stationary points are $(0,0)$, $(0,-1)$, $(0,1)$; the last two are minima, $(0,0)$ is a saddle. Start gradient descent at $\x_0=(1,0)$. The gradient there is $(1,0)$, so the $y$-coordinate stays exactly 0 forever, and the iterates converge to the saddle $(0,0)$. The method did what the theorems promise: it found a stationary point.

The angle between $\d_k$ and $-\grad f$

A line-search method moves $\x_{k+1}=\x_k+\alpha_k\d_k$ along a descent direction $\d_k$, meaning $\g_k^\top\d_k\lt0$ where $\g_k=\grad f(\x_k)$. How "downhill" it is, is measured by the angle $\theta_k$ between $\d_k$ and the steepest-descent direction $-\g_k$:

$$\cos\theta_k=\frac{-\g_k^\top\d_k}{\norm{\g_k}\,\norm{\d_k}}\in(0,1].$$

$\cos\theta_k=1$: $\d_k$ is the steepest-descent direction. $\cos\theta_k\to0$: $\d_k$ becomes perpendicular to the gradient, i.e. it runs along the contour, and moving along it barely changes $f$.

Suppose

  1. $f$ is bounded below on $\R^n$;
  2. $f$ is continuously differentiable on an open set $\mathcal N$ containing the sublevel set $\mathcal L=\{\x:f(\x)\le f(\x_0)\}$, and $\grad f$ is $L$-Lipschitz on $\mathcal N$;
  3. each $\d_k$ is a descent direction and each $\alpha_k$ satisfies the Wolfe conditions ($0\lt c_1\lt c_2\lt1$): $f(\x_k+\alpha_k\d_k)\le f(\x_k)+c_1\alpha_k\g_k^\top\d_k$ and $\grad f(\x_k+\alpha_k\d_k)^\top\d_k\ge c_2\,\g_k^\top\d_k$.

Then $$\sum_{k=0}^\infty\cos^2\theta_k\,\norm{\g_k}^2\lt\infty\qquad\text{(the Zoutendijk condition)}.$$

If in addition $\cos\theta_k\ge\delta>0$ for all $k$ (the directions stay at least a fixed angle away from perpendicular), then $\displaystyle\lim_{k\to\infty}\norm{\g_k}=0$.

Proof. $\delta^2\sum_k\norm{\g_k}^2\le\sum_k\cos^2\theta_k\norm{\g_k}^2\lt\infty$. The terms of a convergent series tend to 0, so $\norm{\g_k}^2\to0$. In particular steepest descent ($\cos\theta_k=1$) with Wolfe steps is globally convergent.

Prove Zoutendijk's theorem. (Examinable: know the hypotheses and the three steps.)

  1. Step 1: the curvature condition forces a step that is not too short. Subtract $\g_k^\top\d_k$ from both sides of the curvature condition: $$(\g_{k+1}-\g_k)^\top\d_k\ge(c_2-1)\,\g_k^\top\d_k.$$ By Cauchy–Schwarz and the Lipschitz property, with $\x_{k+1}-\x_k=\alpha_k\d_k$: $$(\g_{k+1}-\g_k)^\top\d_k\le\norm{\g_{k+1}-\g_k}\norm{\d_k}\le L\alpha_k\norm{\d_k}^2.$$ Chaining the two: $\alpha_k\ge\dfrac{1-c_2}{L}\cdot\dfrac{-\g_k^\top\d_k}{\norm{\d_k}^2}$.

    The curvature condition says "the slope has flattened enough"; Lipschitz continuity says "the slope can't change quickly". Together: you must have travelled some distance. Both $1-c_2$ and $-\g_k^\top\d_k$ are positive, so the bound is a positive number.

  2. Step 2: the Armijo condition turns a long step into a big decrease. Since $\g_k^\top\d_k\lt0$, a lower bound on $\alpha_k$ gives $$f(\x_{k+1})\le f(\x_k)+c_1\alpha_k\g_k^\top\d_k\le f(\x_k)-\frac{c_1(1-c_2)}{L}\frac{(\g_k^\top\d_k)^2}{\norm{\d_k}^2}=f(\x_k)-c\cos^2\theta_k\norm{\g_k}^2,$$ with $c=c_1(1-c_2)/L>0$.

    By the definition of $\theta_k$, $(\g_k^\top\d_k)^2/\norm{\d_k}^2=\cos^2\theta_k\norm{\g_k}^2$. This per-step decrease inequality is the heart of every proof in this chapter.

  3. Step 3: telescope against the lower bound. Summing for $k=0,\dots,N-1$: $$c\sum_{k=0}^{N-1}\cos^2\theta_k\norm{\g_k}^2\le f(\x_0)-f(\x_N)\le f(\x_0)-f^\star\lt\infty.$$ The partial sums of a series of nonnegative terms are bounded, so the series converges.

    The total decrease can never exceed the total "height" $f(\x_0)-f^\star$ available, which is finite because $f$ is bounded below. The iterates stay in $\mathcal L$ because each step decreases $f$, which is why Lipschitz continuity is only needed near $\mathcal L$.

Which directions satisfy the angle condition?

  • Steepest descent $\d_k=-\g_k$: $\cos\theta_k=1$.
  • Newton-like directions $\d_k=-B_k^{-1}\g_k$ with $B_k\succ0$ ([NW] §3.2, p.45). With eigenvalues $\lambda_{\min}\le\lambda_{\max}$ of $B_k$: $-\g_k^\top\d_k=\g_k^\top B_k^{-1}\g_k\ge\norm{\g_k}^2/\lambda_{\max}$ and $\norm{\d_k}\le\norm{\g_k}/\lambda_{\min}$, so $$\cos\theta_k\ge\frac{\lambda_{\min}}{\lambda_{\max}}=\frac1{\kappa(B_k)}.$$ If the condition numbers $\kappa(B_k)$ stay bounded by $M$, then $\cos\theta_k\ge1/M$ and the method is globally convergent. Quasi-Newton methods (Part 10) are analysed this way.

What goes wrong when $\cos\theta_k\to0$? The guaranteed decrease $c\cos^2\theta_k\norm{\g_k}^2$ shrinks to nothing even while the gradient is large. If $\sum_k\cos^2\theta_k\lt\infty$, the total progress can be finite and the method can stall at a point that isn't stationary.

Try it

Each step rotates $-\grad f$ by an angle $\theta_k$ (green arrow: $-\grad f$; blue: the direction used) and then does an exact line search. In "fixed angle" mode, try $\theta=0^\circ$, $60^\circ$, $85^\circ$: the method always converges, just more slowly. Then switch to $\cos\theta_k=\delta_0/(k+1)$ with $\kappa=1$ and run 40 steps: $f$ stops decreasing well above 0, exactly at the predicted level.

Go deeper: the exact stalling level on the round bowl, and Fletcher's version of the theorem

On $f(\x)=\tfrac12\norm{\x}^2$ the gradient is $\x$ itself. An exact line search along a direction at angle $\theta$ to $-\x$ lands at the foot of the perpendicular from the origin onto that line, whose distance from the origin is $\norm{\x}\sin\theta$. So $f_{k+1}=\sin^2\theta_k\,f_k=(1-\cos^2\theta_k)f_k$ and $$f_\infty=f_0\prod_{k\ge0}\Big(1-\frac{\delta_0^2}{(k+1)^2}\Big)=f_0\,\frac{\sin(\pi\delta_0)}{\pi\delta_0}$$ by Euler's product formula for the sine. For $\delta_0=0.5$ the method stalls at $f_\infty=\frac2\pi f_0\approx0.637f_0$, while the gradient there is far from zero. With a constant angle, $f_k=(1-\cos^2\theta)^kf_0\to0$.

[FR] Theorem 2.5.1 (p.30) states the same idea with Fletcher's angle criterion $\theta_k\le\frac\pi2-\mu$ for a fixed $\mu>0$ and Goldstein or Wolfe–Powell line searches: if $\grad f$ is uniformly continuous on the level set, then either $\g_k=\0$ for some $k$, or $f_k\to-\infty$, or $\g_k\to\0$. Fletcher also remarks that steepest descent satisfies the angle criterion trivially yet is often slow, which is why rates (the rest of this part) matter as much as convergence.

Constant step sizes ([Y] §1.2.3)

Nesterov studies the plainest scheme: gradient descent $\x_{k+1}=\x_k-h_k\grad f(\x_k)$ with step $h_k$. In the Nesterov-based chapters the step is called $h$ (Part 5 called it $\alpha$). He compares three rules: a constant step $h_k=h$; full relaxation (exact line search) $h_k=\argmin_{h\ge0}f(\x_k-h\grad f(\x_k))$; and the Goldstein–Armijo rule, which with $\x_k-\x_{k+1}=h_k\g_k$ asks for $$a\,h_k\norm{\g_k}^2\le f(\x_k)-f(\x_{k+1})\le b\,h_k\norm{\g_k}^2,\qquad0\lt a\lt b\lt1$$ ([Y] writes $\alpha,\beta$ for $a,b$).

Assume $f\in C_L^{1,1}$ (gradient $L$-Lipschitz; no convexity) and bounded below by $f^\star$. The Descent Lemma (Part 6) with $\y=\x_k-h\g_k$ gives [Y] (1.2.12): $$f(\x_{k+1})\le f(\x_k)-h\norm{\g_k}^2+\tfrac{h^2L}2\norm{\g_k}^2=f(\x_k)-h\Big(1-\tfrac h2L\Big)\norm{\g_k}^2.$$

  • Constant step $h$: $f(\x_k)-f(\x_{k+1})\ge h(1-\tfrac12hL)\norm{\g_k}^2$, positive exactly when $0\lt h\lt2/L$. The guaranteed decrease $\Delta(h)=h(1-\tfrac12hL)$ is largest at $h^\ast=1/L$, giving $\frac1{2L}\norm{\g_k}^2$.
  • Full relaxation: $f(\x_k)-f(\x_{k+1})\ge\frac1{2L}\norm{\g_k}^2$ (the exact minimizer does at least as well as $h=1/L$).
  • Goldstein–Armijo: the upper inequality and (1.2.12) give $b\,h_k\ge h_k(1-\frac{h_k}2L)$, so $h_k\ge\frac2L(1-b)$; the lower inequality then gives $f(\x_k)-f(\x_{k+1})\ge\frac2La(1-b)\norm{\g_k}^2$.

In all cases $f(\x_k)-f(\x_{k+1})\ge\dfrac\omega L\norm{\g_k}^2$ for a constant $\omega>0$ ([Y] (1.2.13)).

Summing (1.2.13) for $k=0,\dots,N$: $\displaystyle\frac\omega L\sum_{k=0}^N\norm{\g_k}^2\le f(\x_0)-f^\star$. Hence $\norm{\g_k}\to0$, and with $g_N^\ast=\min_{0\le k\le N}\norm{\g_k}$, $$g_N^\ast\le\frac1{\sqrt{N+1}}\Big[\frac L\omega\big(f(\x_0)-f^\star\big)\Big]^{1/2}.$$ For a constant step, $\omega/L=h(1-\frac12hL)$; at $h=1/L$ this reads $g_N^\ast\le\sqrt{2L(f(\x_0)-f^\star)/(N+1)}$.

Careful: this is a rate for the best gradient seen so far, of order $1/\sqrt N$. It says nothing about the rate of $f(\x_k)$ or of $\x_k$ ([Y] p.28), and the limit may be a saddle. Chapters 7.3–7.4 show what convexity adds.
Try it

$f(x,y)=\tfrac12x^2+2\cos y$ is not convex, has $L=2$ and $f^\star=-2$. Compare $hL=1$ (best guarantee), $hL=0.2$ (tiny steps, weak guarantee) and $hL=2.2$ (beyond $2/L$: no guarantee; watch the path). Then press "Start on the ridge": like Nesterov's example, the method converges to a saddle. The lower plot shows that the observed $\min\norm{\grad f}^2$ always sits below the guaranteed curve when $hL\lt2$.

For $f(x,y)=\tfrac12x^2+2\cos y$, find $L$ and $f^\star$, and use (1.2.15) to bound $\min_{0\le k\le99}\norm{\grad f(\x_k)}$ for $h=1/L$ from $\x_0=(2,0.5)$.

  1. $\grad f=(x,-2\sin y)$, $\hess f=\mathrm{diag}(1,-2\cos y)$. Its eigenvalues lie in $[-2,2]$, so $\norm{\hess f}\le2$ and $L=2$.

    For $C^2$ functions, $\grad f$ is $L$-Lipschitz when every Hessian eigenvalue has absolute value at most $L$ (Part 4/6). Negative curvature counts too.

  2. $f\ge0+2(-1)=-2$, attained at $(0,\pi)$: $f^\star=-2$. And $f(\x_0)=2+2\cos0.5=3.7552$, so $f(\x_0)-f^\star=5.7552$.

    The bound needs the total available decrease.

  3. $h=1/L=0.5$, $N=99$: $g_{99}^\ast\le\sqrt{2\cdot2\cdot5.7552/100}=\sqrt{0.2302}=0.480$.

    A worst-case guarantee is usually pessimistic: in the widget the actual gradient is tiny long before 100 steps. Its value is that it holds for every $L$-smooth function bounded below.

At $\x_k$ the gradient is $\g_k=(2,1)$ and the method uses $\d_k=(-1,1)$. Is $\d_k$ a descent direction, and what is $\cos\theta_k$?

$\cos\theta_k=-\g_k^\top\d_k/(\norm{\g_k}\norm{\d_k})$. A positive value means descent.

$-\g_k^\top\d_k=-(-2+1)=1>0$, so yes. $\norm{\g_k}=\sqrt5$, $\norm{\d_k}=\sqrt2$, so $\cos\theta_k=1/\sqrt{10}\approx0.316$ ($\theta_k\approx71.6^\circ$).

A Newton-like method uses $\d_k=-B_k^{-1}\g_k$ with eigenvalues of $B_k$ always in $[1,100]$. What lower bound on $\cos\theta_k$ follows, valid for every $k$?

$\cos\theta_k\ge\lambda_{\min}(B_k)/\lambda_{\max}(B_k)$.

$\cos\theta_k\ge1/100=0.01$. Since this is a fixed positive $\delta$, Zoutendijk's corollary gives $\norm{\g_k}\to0$ with Wolfe steps.

$f$ is $L$-smooth with $L=4$. Gradient descent uses the constant step $h=0.4$, and at the current point $\norm{\g_k}=3$. What decrease $f(\x_k)-f(\x_{k+1})$ is guaranteed?

$h(1-\tfrac12hL)\norm{\g_k}^2$.

$h(1-\tfrac12hL)=0.4(1-0.8)=0.08$, times $\norm{\g_k}^2=9$: $0.72$. (With the best step $h=1/L=0.25$ the guarantee would be $\frac1{2L}\cdot9=1.125$.)

$f$ is $L$-smooth with $L=2$, bounded below, and $f(\x_0)-f^\star=5$. With $h=1/L$, what is the smallest $N$ for which (1.2.15) guarantees $\min_{0\le k\le N}\norm{\g_k}\le0.01$?

Require $\sqrt{2L(f(\x_0)-f^\star)/(N+1)}\le0.01$ and solve for $N$.

$N+1\ge2L(f(\x_0)-f^\star)/\varepsilon^2=20/10^{-4}=200000$, so $N=199999$. The $1/\sqrt N$ rate makes small gradient tolerances expensive: halving $\varepsilon$ quadruples $N$.

  • State all three Zoutendijk hypotheses: bounded below, Lipschitz gradient near the sublevel set, Wolfe steps along descent directions
  • Reproduce the skeleton: per-step decrease, telescope against $f(\x_0)-f^\star$, conclude
  • Check the angle condition $\cos\theta_k\ge\delta>0$ for any new direction rule
  • For constant steps, keep $0\lt h\lt2/L$; $h=1/L$ gives the best guarantee
  • Saying "globally convergent" means "converges to the global minimizer"
  • Forgetting that the Zoutendijk sum alone gives $\cos^2\theta_k\norm{\g_k}^2\to0$, not $\norm{\g_k}\to0$, unless $\cos\theta_k$ is bounded away from 0
  • Reading (1.2.15) as a rate for $f(\x_k)$ or for $\x_k$
  1. Zoutendijk: bounded below + Lipschitz gradient + Wolfe steps $\Rightarrow\sum_k\cos^2\theta_k\norm{\g_k}^2\lt\infty$; with $\cos\theta_k\ge\delta>0$ this gives $\norm{\g_k}\to0$.
  2. For $L$-smooth $f$, each constant step ($0\lt h\lt2/L$), exact step, or Goldstein–Armijo step decreases $f$ by at least $\frac\omega L\norm{\g_k}^2$; $h=1/L$ is the best constant step.
  3. Without convexity: only convergence to a stationary point, with $\min_{k\le N}\norm{\g_k}=O(1/\sqrt N)$.

Zoutendijk's theorem concludes, for a method with Wolfe steps, that…

$\x_k$ converges to a global minimizer
The theorem uses no convexity at all. What does it say about gradients?
$\sum_k\cos^2\theta_k\norm{\grad f(\x_k)}^2$ is finite
That's the Zoutendijk condition. Convergence of $\norm{\g_k}$ needs the extra angle condition.
$\norm{\grad f(\x_k)}\to0$ for every choice of descent directions
Directions that turn perpendicular to the gradient can stall the method, as the Try-it showed.

In the proof, which hypothesis supplies the lower bound $\alpha_k\ge\frac{1-c_2}L\frac{-\g_k^\top\d_k}{\norm{\d_k}^2}$?

The Armijo condition with bounded-below $f$
Armijo is used next, to turn the step bound into a decrease. Which condition stops steps from being too short?
The curvature condition together with the Lipschitz gradient
Curvature: $(\g_{k+1}-\g_k)^\top\d_k\ge(c_2-1)\g_k^\top\d_k$; Lipschitz: the same quantity is $\le L\alpha_k\norm{\d_k}^2$.
Boundedness below alone
That's used in the telescoping step, to bound the total decrease.

For an $L$-smooth $f$, the constant step $h=1.5/L$…

is unsafe because it exceeds $1/L$
Evaluate $h(1-\frac12hL)$ at $h=1.5/L$.
still guarantees decrease, but only $\frac{0.375}{L}\norm{\g_k}^2$ per step
$h(1-\frac12hL)=\frac{1.5}L(1-0.75)=\frac{0.375}L$, positive since $h\lt2/L$, but smaller than the best value $\frac{0.5}L$ at $h=1/L$.
guarantees more decrease than $h=1/L$ because the step is longer
The guaranteed decrease $h(1-\frac12hL)$ is a downward parabola in $h$ with its peak at $1/L$.

Gradient descent with $h=1/L$ is started at $(1,0)$ on Nesterov's $f=\tfrac12x^2+\tfrac14y^4-\tfrac12y^2$. Which statement is consistent with the theory?

It must converge to a minimizer because $f$ is bounded below
Bounded below only gives $\grad f\to0$.
It cannot converge because the gradient isn't globally Lipschitz
Along the iterates' path $y=0$ the function is just $\tfrac12x^2$; the method converges, but to what?
It converges to the saddle $(0,0)$, a stationary point that is not a minimizer
The $y$-gradient is $y^3-y=0$ at $y=0$, so $y$ never moves and $x\to0$.

Convexity buys a rate: $O(1/k)$ for $\mathcal F_L^{1,1}$

For convex functions with $L$-Lipschitz gradient, gradient descent with constant step drives the optimality gap $f(\x_k)-f^\star$ to zero at least as fast as $C/k$, at every iterate.

This is the first rate in the course for the function value itself (not just the best gradient so far), and its four-step proof (bounded iterates, decrease, convexity, inverting a recursion) is a model you'll be asked to reproduce.

A smoke detector that beeps louder the more smoke there is: convexity makes the gradient "beep" in proportion to how far $f$ is from optimal, so a large gap can't hide behind a small gradient.

Setting: $f\in\mathcal F_L^{1,1}(\R^n)$ (convex, differentiable, $\grad f$ $L$-Lipschitz; Part 4), with a minimizer $\x^\star$ (assumed to exist), $f^\star=f(\x^\star)$. Method: $\x_{k+1}=\x_k-h\grad f(\x_k)$ with constant $0\lt h\lt2/L$. Notation: $r_k=\norm{\x_k-\x^\star}$ and $\Delta_k=f(\x_k)-f^\star$.

The idea in one line

Convexity gives $f^\star\ge f(\x_k)+\g_k^\top(\x^\star-\x_k)$, i.e. $\Delta_k\le\g_k^\top(\x_k-\x^\star)\le\norm{\g_k}\,r_k$. So a large gap forces a large gradient, and (Descent Lemma) a large gradient forces a large decrease. This gives a self-improving recursion $$\Delta_{k+1}\le\Delta_k-c\,\Delta_k^2,$$ and any such recursion decays like $1/(ck)$. Try it with $c=1$, $\Delta_0=0.5$: $0.5,\ 0.25,\ 0.1875,\ 0.1523,\ 0.1291,\ 0.1125,\dots$ (compare $1/(k+2)$: $0.5,0.333,0.25,0.2,0.167,0.143$).

One more ingredient is needed: $r_k$ must not grow. That comes from a sharper inequality than convexity.

If $f\in\mathcal F_L^{1,1}$, then for all $\x,\y$: $$f(\y)\ge f(\x)+\grad f(\x)^\top(\y-\x)+\frac1{2L}\norm{\grad f(\y)-\grad f(\x)}^2,\tag{$\star$}$$ $$\big(\grad f(\x)-\grad f(\y)\big)^\top(\x-\y)\ge\frac1L\norm{\grad f(\x)-\grad f(\y)}^2.\tag{co-coercivity}$$

Proof. Fix $\x$ and let $\varphi(\z)=f(\z)-\grad f(\x)^\top\z$. It is convex and $L$-smooth (we only subtracted a linear function), and $\grad\varphi(\x)=\0$, so $\x$ is a global minimizer of $\varphi$ (convex + zero gradient, Part 4). Apply the Descent Lemma to $\varphi$ at the point $\y$ with the step $\z=\y-\frac1L\grad\varphi(\y)$: $$\varphi(\x)\le\varphi(\z)\le\varphi(\y)-\tfrac1L\norm{\grad\varphi(\y)}^2+\tfrac L2\cdot\tfrac1{L^2}\norm{\grad\varphi(\y)}^2=\varphi(\y)-\tfrac1{2L}\norm{\grad\varphi(\y)}^2.$$ Since $\grad\varphi(\y)=\grad f(\y)-\grad f(\x)$ and $\varphi(\y)-\varphi(\x)=f(\y)-f(\x)-\grad f(\x)^\top(\y-\x)$, this is ($\star$). Write ($\star$) again with $\x$ and $\y$ swapped and add the two: the function values cancel and co-coercivity remains. $\square$

Compared with plain convexity, ($\star$) has the extra term $\frac1{2L}\norm{\grad f(\y)-\grad f(\x)}^2$: convexity and smoothness working together.

Let $f\in\mathcal F_L^{1,1}(\R^n)$ and $0\lt h\lt2/L$. Then gradient descent with constant step $h$ satisfies $$f(\x_k)-f^\star\le\frac{2\big(f(\x_0)-f^\star\big)\norm{\x_0-\x^\star}^2}{2\norm{\x_0-\x^\star}^2+k\,h(2-Lh)\big(f(\x_0)-f^\star\big)}.$$

With $h=1/L$: $$f(\x_k)-f^\star\le\frac{2L\norm{\x_0-\x^\star}^2}{k+4}.$$ To guarantee $f(\x_k)-f^\star\le\varepsilon$ it suffices that $k\ge\dfrac{2L\norm{\x_0-\x^\star}^2}{\varepsilon}-4$.

Prove Theorem 2.1.14 and Corollary 2.1.2. (Examinable: know the rate by heart and be able to give the steps.)

  1. Step 1: the iterates never move away from $\x^\star$. Expand $$r_{k+1}^2=\norm{\x_k-\x^\star-h\g_k}^2=r_k^2-2h\,\g_k^\top(\x_k-\x^\star)+h^2\norm{\g_k}^2.$$ Co-coercivity with $\y=\x^\star$ and $\grad f(\x^\star)=\0$ gives $\g_k^\top(\x_k-\x^\star)\ge\frac1L\norm{\g_k}^2$, so $$r_{k+1}^2\le r_k^2-h\Big(\frac2L-h\Big)\norm{\g_k}^2\le r_k^2.$$ Hence $r_k\le r_0$ for all $k$.

    This is where $h\lt2/L$ and co-coercivity enter. Without it we'd have no control over $r_k$ in Step 3.

  2. Step 2: per-step decrease. The Descent Lemma (as in Chapter 7.2): $$f(\x_{k+1})\le f(\x_k)-\omega\norm{\g_k}^2,\qquad\omega=h\Big(1-\frac L2h\Big)=\tfrac12h(2-Lh)>0.$$

    Same inequality as in the non-convex case. Convexity hasn't been used yet for this step.

  3. Step 3: convexity turns the gap into a gradient. By the first-order convexity inequality and Cauchy–Schwarz, $$\Delta_k\le\g_k^\top(\x_k-\x^\star)\le\norm{\g_k}\,r_k\le r_0\norm{\g_k},$$ so $\norm{\g_k}^2\ge\Delta_k^2/r_0^2$, and Step 2 becomes $\Delta_{k+1}\le\Delta_k-\dfrac\omega{r_0^2}\Delta_k^2$.

    Here Step 1 is essential: we need one fixed constant $r_0$ for every $k$.

  4. Step 4: invert and telescope. If some $\Delta_k=0$ we are done. Otherwise divide by $\Delta_k\Delta_{k+1}>0$ and use $\Delta_{k+1}\le\Delta_k$: $$\frac1{\Delta_{k+1}}\ge\frac1{\Delta_k}+\frac\omega{r_0^2}\cdot\frac{\Delta_k}{\Delta_{k+1}}\ge\frac1{\Delta_k}+\frac\omega{r_0^2}.$$ Summing from $0$ to $k-1$: $\dfrac1{\Delta_k}\ge\dfrac1{\Delta_0}+\dfrac{k\omega}{r_0^2}$, i.e. $\Delta_k\le\dfrac{\Delta_0r_0^2}{r_0^2+k\omega\Delta_0}$. Multiply top and bottom by 2 and use $2\omega=h(2-Lh)$: that's the theorem.

    The reciprocal $1/\Delta_k$ grows at least linearly in $k$. That is exactly a $1/k$ rate for $\Delta_k$.

  5. Corollary ($h=1/L$). Now $h(2-Lh)=1/L$ and the bound is $\dfrac{2\Delta_0r_0^2}{2r_0^2+k\Delta_0/L}$. As a function of $\Delta_0$ this is increasing (it has the form $\frac{a\Delta_0}{b+c\Delta_0}$ with $a,b,c>0$). The Descent Lemma at $\x^\star$ (where $\grad f(\x^\star)=\0$) gives $\Delta_0\le\frac L2r_0^2$. Substituting this largest possible $\Delta_0$: $$\Delta_k\le\frac{2\cdot\frac L2r_0^2\cdot r_0^2}{2r_0^2+\frac k2r_0^2}=\frac{Lr_0^2}{2+k/2}=\frac{2Lr_0^2}{k+4}.$$

    Replacing $\Delta_0$ by an upper bound is only allowed because the bound is increasing in $\Delta_0$ ([Y] p.70 states this explicitly).

Try it

This is the book's lab: constant-step gradient descent with $h=1/L$, $L=1$, on a 500-dimensional diagonal quadratic from $\x_0=(1,\dots,1)$. With $\lambda_i=1/i^2$ the smallest eigenvalue is $4\times10^{-6}$: convex, but effectively not strongly convex. On log–log axes the gap is a straight line, sitting under the dashed $O(1/k)$ bound. Then switch to the second problem and to semi-log axes to see the contrast that Chapter 7.4 explains.

Go deeper: the version from the lecture, for any descent direction and Goldstein–Armijo steps

The class diary records that the 3 September proof "holds for any descent direction and Armijo–Goldstein" and was then specialized to the constant step along $-\grad f$ (which is Theorem 2.1.14). The 2024 class diary says the same: in class the analysis was done "for any descent direction with inexact line search", there is "no good reference", and notes were handed out. The textbooks and the lecture-notes book only write out the constant-step case, so here is how the same skeleton extends; the handout's exact assumptions may differ in detail, so check it against your class notes. [Y] also remarks (p.69) that the rate for other reasonable step-size rules is similar.

Let $\x_{k+1}=\x_k+t_k\d_k$ with $\cos\theta_k\ge\delta>0$ and $t_k$ satisfying Goldstein–Armijo along $\d_k$: $a\,t_k(-\g_k^\top\d_k)\le f(\x_k)-f(\x_{k+1})\le b\,t_k(-\g_k^\top\d_k)$, $0\lt a\lt b\lt1$.

  1. Decrease. The Descent Lemma gives $f(\x_k)-f(\x_{k+1})\ge t_k(-\g_k^\top\d_k)-\frac L2t_k^2\norm{\d_k}^2$. Combined with the upper Goldstein inequality: $t_k\ge\frac{2(1-b)}L\frac{-\g_k^\top\d_k}{\norm{\d_k}^2}$. The lower inequality then gives $f(\x_k)-f(\x_{k+1})\ge\frac{2a(1-b)}L\cos^2\theta_k\norm{\g_k}^2\ge c\norm{\g_k}^2$ with $c=\frac{2a(1-b)\delta^2}L$.
  2. Boundedness. Co-coercivity no longer controls $r_k$ for a general direction. Instead, $f$ decreases, so all iterates lie in $\{\x:f(\x)\le f(\x_0)\}$; assume this set is bounded, with $R=\max\norm{\x-\x^\star}$ over it.
  3. Convexity gives $\Delta_k\le R\norm{\g_k}$, hence $\Delta_{k+1}\le\Delta_k-\frac c{R^2}\Delta_k^2$, and inverting as in Step 4: $\frac1{\Delta_k}\ge\frac1{\Delta_0}+\frac{ck}{R^2}$, an $O(1/k)$ rate.

A related exercise in the book's problem appendix lets the constant step vary, $h_k\in(0,2/L)$: then $f(\x_N)-f^\star\le2r_0^2\big/\sum_{k\lt N}h_k(2-Lh_k)$.

Go deeper: is $1/k$ the best possible?

For gradient descent with constant step on $\mathcal F_L^{1,1}$, the $1/k$ order cannot be improved in the worst case. But among all first-order methods (methods that only use gradients), Nesterov's lower complexity bound for this class is of order $1/k^2$ ([Y] Thm 2.1.7), and his accelerated gradient method ([Y] §2.2) achieves it. That is beyond this course; just don't claim that $O(1/k)$ is optimal for every first-order method.

$f\in\mathcal F_L^{1,1}$ with $L=2$ and $\norm{\x_0-\x^\star}=3$. With $h=1/L$, bound $f(\x_{96})-f^\star$, and find how many iterations guarantee $f(\x_k)-f^\star\le10^{-3}$.

  1. $f(\x_{96})-f^\star\le\dfrac{2\cdot2\cdot9}{96+4}=0.36$.

    Direct substitution into Corollary 2.1.2.

  2. $\dfrac{36}{k+4}\le10^{-3}\iff k\ge36000-4=35996$.

    Each extra digit of accuracy costs ten times as many iterations: the price of a sublinear rate.

$f\in\mathcal F_L^{1,1}$ with $L=10$ and $\norm{\x_0-\x^\star}=1$. Using Corollary 2.1.2 ($h=1/L$), what is the smallest $k$ that guarantees $f(\x_k)-f^\star\le10^{-3}$?

Solve $2Lr_0^2/(k+4)\le\varepsilon$.

$k+4\ge2\cdot10\cdot1/10^{-3}=20000$, so $k=19996$.

Use the full Theorem 2.1.14 with $L=1$, $h=1$, $\norm{\x_0-\x^\star}=2$ and $f(\x_0)-f^\star=1$. What bound does it give on $f(\x_{10})-f^\star$?

$h(2-Lh)=1$. Plug into $\dfrac{2\Delta_0r_0^2}{2r_0^2+kh(2-Lh)\Delta_0}$.

$\dfrac{2\cdot1\cdot4}{2\cdot4+10\cdot1\cdot1}=\dfrac8{18}=\dfrac49\approx0.444$. Corollary 2.1.2 gives the weaker $\frac{2\cdot1\cdot4}{14}\approx0.571$, because it replaces $\Delta_0=1$ by its worst case $\frac L2r_0^2=2$.

$f\in\mathcal F_L^{1,1}$ with $L=4$ and $\norm{\x_0-\x^\star}=2$. What is the largest possible value of $f(\x_0)-f^\star$?

Apply the Descent Lemma at $\x^\star$, where the gradient is zero.

$f(\x_0)\le f^\star+\grad f(\x^\star)^\top(\x_0-\x^\star)+\frac L2\norm{\x_0-\x^\star}^2=f^\star+\frac42\cdot4$, so the gap is at most 8, attained by $f(\x)=2\norm{\x-\x^\star}^2$.

A sequence satisfies $\Delta_{k+1}\le\Delta_k-\frac14\Delta_k^2$ with $\Delta_0=1$ and $\Delta_k>0$ decreasing. What upper bound on $\Delta_{20}$ does the inversion trick give?

Step 4 gives $\frac1{\Delta_k}\ge\frac1{\Delta_0}+\frac k4$.

$\frac1{\Delta_{20}}\ge1+5=6$, so $\Delta_{20}\le\frac16\approx0.1667$.

  • Name the four steps: $r_k\le r_0$ (co-coercivity), per-step decrease (Descent Lemma), $\Delta_k\le r_0\norm{\g_k}$ (convexity + Cauchy–Schwarz), invert and telescope
  • Remember the corollary: $f(\x_k)-f^\star\le2L\norm{\x_0-\x^\star}^2/(k+4)$ at $h=1/L$
  • Check that a minimizer exists before quoting the theorem
  • Confusing this sublinear $O(1/k)$ rate with the linear rate of strongly convex functions
  • Dividing the recursion by $\Delta_k\Delta_{k+1}$ without noting $\Delta_{k+1}\le\Delta_k$
  • Substituting an upper bound for $\Delta_0$ without checking that the bound is increasing in $\Delta_0$
  1. Co-coercivity: $(\grad f(\x)-\grad f(\y))^\top(\x-\y)\ge\frac1L\norm{\grad f(\x)-\grad f(\y)}^2$ for $f\in\mathcal F_L^{1,1}$.
  2. For $f\in\mathcal F_L^{1,1}$ and $0\lt h\lt2/L$: $f(\x_k)-f^\star=O(1/k)$ at every iterate; at $h=1/L$, $\le2L\norm{\x_0-\x^\star}^2/(k+4)$.
  3. The engine is the recursion $\Delta_{k+1}\le\Delta_k-c\Delta_k^2$, which after inverting gives $1/\Delta_k\ge1/\Delta_0+ck$.

In the proof of Theorem 2.1.14, where is convexity used?

Only in the Descent Lemma step
The Descent Lemma holds for every $L$-smooth function, convex or not.
In co-coercivity (to keep $r_k\le r_0$) and in $\Delta_k\le\g_k^\top(\x_k-\x^\star)$
Both need convexity: co-coercivity is proved via "convex + zero gradient ⇒ minimizer", and the gap bound is the tangent-plane inequality.
Only to guarantee that a minimizer exists
Convex functions need not have minimizers ($e^x$); existence is an extra assumption here.

On a log–log plot of $f(\x_k)-f^\star$ against $k$, an $O(1/k)$ method looks like…

a straight line of slope about $-1$
$\log(C/k)=\log C-\log k$.
a straight line on semi-log axes
That's the signature of geometric convergence.
a curve that plunges ever more steeply
That's superlinear behaviour.

Why can't the $O(1/k)$ argument be run for a non-convex $L$-smooth $f$?

Because the Descent Lemma fails without convexity
It doesn't fail; Chapter 7.2 used it on a non-convex function.
Because $h=1/L$ is not allowed
The step rule is the same; something in the inequalities breaks.
Because a small gradient no longer forces a small gap: $\Delta_k\le r_0\norm{\g_k}$ fails
At a saddle the gradient is zero while $f-f^\star$ can be large; only convexity links them.

Which bound is Nesterov's Corollary 2.1.2 for $h=1/L$?

$f(\x_k)-f^\star\le\frac{L\norm{\x_0-\x^\star}^2}{2k}$
Close in spirit, but not the stated form. Recall the "+4".
$f(\x_k)-f^\star\le\big(\frac{L-1}{L+1}\big)^k$
A geometric rate needs strong convexity.
$f(\x_k)-f^\star\le\frac{2L\norm{\x_0-\x^\star}^2}{k+4}$
Obtained from the theorem by bounding $\Delta_0\le\frac L2r_0^2$.

Strong convexity buys a linear rate: Theorem 2.1.15

If $f$ is $\mu$-strongly convex and $L$-smooth, gradient descent with step $h=2/(\mu+L)$ shrinks the distance to the minimizer by the factor $\frac{Q_f-1}{Q_f+1}$ at every step, where $Q_f=L/\mu$ is the condition number.

It explains in general what Part 5 found for quadratics: the condition number sets the speed, and the iteration count is about $\frac{Q_f}2\ln\frac1\varepsilon$. It's the capstone result of the midsem syllabus.

Draining a tank through a valve: the outflow is proportional to the water left, so a fixed fraction leaves every minute, never slowing down in relative terms. Strong convexity makes the "pull" towards $\x^\star$ proportional to the distance.

Recall (Part 4): $f\in\mathcal S_{\mu,L}^{1,1}$ if $f\in\mathcal F_L^{1,1}$ and $f$ is $\mu$-strongly convex, $f(\y)\ge f(\x)+\grad f(\x)^\top(\y-\x)+\frac\mu2\norm{\y-\x}^2$; for $C^2$ functions, $\mu I\preceq\hess f(\x)\preceq LI$. Such $f$ has exactly one minimizer $\x^\star$. Nesterov writes $Q_f=L/\mu\ge1$ for the condition number; it plays the role of $\kappa$ from Part 5 (for a quadratic, $Q_f=\kappa$).

If $f\in\mathcal S_{\mu,L}^{1,1}$, then for all $\x,\y$: $$\big(\grad f(\x)-\grad f(\y)\big)^\top(\x-\y)\ge\frac{\mu L}{\mu+L}\norm{\x-\y}^2+\frac1{\mu+L}\norm{\grad f(\x)-\grad f(\y)}^2.$$

Go deeper: proof of strong co-coercivity

If $\mu=L$, $f$ is $\frac L2\norm{\x}^2$ plus an affine function and both sides equal $L\norm{\x-\y}^2$. If $\mu\lt L$, let $g(\z)=f(\z)-\frac\mu2\norm{\z}^2$. For $C^2$ $f$, $0\preceq\hess g=\hess f-\mu I\preceq(L-\mu)I$, so $g\in\mathcal F_{L-\mu}^{1,1}$ and plain co-coercivity applies to $g$. Write $A=(\grad f(\x)-\grad f(\y))^\top(\x-\y)$, $B=\norm{\grad f(\x)-\grad f(\y)}^2$, $C=\norm{\x-\y}^2$. Since $\grad g(\z)=\grad f(\z)-\mu\z$, co-coercivity for $g$ reads $$A-\mu C\ge\frac{B-2\mu A+\mu^2C}{L-\mu}.$$ Multiply by $L-\mu$ and collect: $(L+\mu)A\ge B+\mu LC$. Divide by $L+\mu$. $\square$

Let $f\in\mathcal S_{\mu,L}^{1,1}(\R^n)$ and $0\lt h\le\dfrac2{\mu+L}$. Gradient descent with constant step $h$ satisfies $$\norm{\x_k-\x^\star}^2\le\Big(1-\frac{2h\mu L}{\mu+L}\Big)^k\norm{\x_0-\x^\star}^2.$$ If $h=\dfrac2{\mu+L}$, then with $Q_f=L/\mu$: $$\norm{\x_k-\x^\star}\le\Big(\frac{Q_f-1}{Q_f+1}\Big)^k\norm{\x_0-\x^\star},\qquad f(\x_k)-f^\star\le\frac L2\Big(\frac{Q_f-1}{Q_f+1}\Big)^{2k}\norm{\x_0-\x^\star}^2.$$

Prove Theorem 2.1.15 (the general step $h$, then the special step $2/(\mu+L)$, then the function-value bound).

  1. As in Chapter 7.3, $r_{k+1}^2=r_k^2-2h\,\g_k^\top(\x_k-\x^\star)+h^2\norm{\g_k}^2$.

    Just expanding $\norm{\x_k-h\g_k-\x^\star}^2$.

  2. Strong co-coercivity with $\y=\x^\star$, $\grad f(\x^\star)=\0$: $\g_k^\top(\x_k-\x^\star)\ge\frac{\mu L}{\mu+L}r_k^2+\frac1{\mu+L}\norm{\g_k}^2$. Substituting: $$r_{k+1}^2\le\Big(1-\frac{2h\mu L}{\mu+L}\Big)r_k^2+h\Big(h-\frac2{\mu+L}\Big)\norm{\g_k}^2.$$

    The new $r_k^2$ term (from strong convexity) is what makes the contraction a fixed fraction of $r_k^2$.

  3. For $0\lt h\le\frac2{\mu+L}$ the last coefficient is $\le0$, so $r_{k+1}^2\le\big(1-\frac{2h\mu L}{\mu+L}\big)r_k^2$. Iterating gives the general bound.

    This is the "general case": any step up to $2/(\mu+L)$.

  4. At $h=\frac2{\mu+L}$: $1-\frac{4\mu L}{(\mu+L)^2}=\frac{(L-\mu)^2}{(L+\mu)^2}=\big(\frac{Q_f-1}{Q_f+1}\big)^2$ (divide top and bottom by $\mu^2$). Taking square roots gives the distance bound.

    This is the largest step the theorem allows and it gives the smallest factor.

  5. Function values: the Descent Lemma at $\x^\star$ gives $f(\x_k)\le f^\star+\frac L2r_k^2$, so $f(\x_k)-f^\star\le\frac L2\big(\frac{Q_f-1}{Q_f+1}\big)^{2k}r_0^2$.

    The gap is bounded by a constant times the squared distance, so it converges with the squared rate.

Closing the loop with Part 5

On a quadratic $f(\x)=\tfrac12\x^\top A\x-\b^\top\x$ with eigenvalues in $[m,L]$, Part 5 found that constant step $\alpha$ gives $\x_{k+1}-\x^\star=(I-\alpha A)(\x_k-\x^\star)$, worst-case factor $\max_i|1-\alpha\lambda_i|$, minimized at $\alpha=\frac2{m+L}$ with value $\frac{\kappa-1}{\kappa+1}$. Theorem 2.1.15 gives the same step and the same factor for every $f\in\mathcal S_{\mu,L}^{1,1}$, quadratic or not. And the exact-line-search rate $\big(\frac{\kappa-1}{\kappa+1}\big)^2$ per step in function value (Kantorovich, Part 5) matches the $\big(\frac{Q_f-1}{Q_f+1}\big)^{2k}$ in the $f$-bound here.

Try it

The blue curve is the exact worst-case factor on quadratics with eigenvalues in $[\mu,L]$ (Part 5); the dashed green curve is the theorem's guarantee $\sqrt{1-2h\mu L/(\mu+L)}$ for all of $\mathcal S_{\mu,L}^{1,1}$. Move $h$: they touch at $h=2/(\mu+L)$. Increase $Q_f$ and watch the best factor creep towards 1 and the iteration count explode.

Try it

Back to the lab, now with $\lambda_i$ spread over $[\mu,1]$ and semi-log axes. With $\mu=0.05$ the gap drops from about $0.4$ at $k=10$ to about $3\times10^{-6}$ at $k=100$: a straight line, under the dashed linear bound. Lower $\mu$ towards $10^{-4}$: the straight line flattens and the $O(1/k)$ bound becomes the better description for a long time.

$f(x)=\tfrac12x^2+\ln(1+e^x)$ on $\R$. Find $\mu$, $L$, the step $h=2/(\mu+L)$, the rate, and the number of iterations that guarantees $|x_k-x^\star|\le10^{-6}|x_0-x^\star|$. Then do one step from $x_0=2$.

  1. With $\sigma(x)=1/(1+e^{-x})$: $f'(x)=x+\sigma(x)$ and $f''(x)=1+\sigma(x)(1-\sigma(x))$. Since $\sigma(1-\sigma)\in(0,\tfrac14]$, $f''\in[1,1.25]$: $\mu=1$, $L=1.25$.

    For a $C^2$ function, bounds on the Hessian give $\mu$ and $L$ directly. This $f$ is not a quadratic.

  2. $h=\frac2{2.25}=\frac89$, $Q_f=1.25$, $\frac{Q_f-1}{Q_f+1}=\frac{0.25}{2.25}=\frac19$.

    A well-conditioned function: each step cuts the distance to $x^\star$ by a factor of at least 9.

  3. $(1/9)^k\le10^{-6}\iff k\ge6/\log_{10}9=6.29$, so $k=7$.

    Iterations for a linear rate: $\ln(1/\varepsilon)/\ln(1/\rho)$, rounded up.

  4. $x^\star\approx-0.4011$ (solving $x+\sigma(x)=0$). From $x_0=2$: $x_1=2-\frac89(2+\sigma(2))=2-\frac89(2.8808)=-0.5607$. Then $|x_1-x^\star|=0.160\le\frac19\cdot2.401=0.267$.

    The theorem's bound holds, with room to spare: it is a worst case over the whole class.

Assumptions on $f$StepGuaranteeRate
$L$-smooth, bounded below (Ch. 7.2)$0\lt h\lt2/L$, exact, Goldstein–Armijo, or Wolfe$\min_{k\le N}\norm{\g_k}\le C/\sqrt{N+1}$; $\g_k\to\0$sublinear, stationarity only
$\mathcal F_L^{1,1}$: convex + $L$-smooth (Ch. 7.3)$0\lt h\lt2/L$$f(\x_k)-f^\star\le\frac{2L r_0^2}{k+4}$ at $h=1/L$sublinear $O(1/k)$
$\mathcal S_{\mu,L}^{1,1}$: strongly convex + $L$-smooth (Ch. 7.4)$0\lt h\le\frac2{\mu+L}$$\norm{\x_k-\x^\star}\le\big(\frac{Q_f-1}{Q_f+1}\big)^kr_0$ at $h=\frac2{\mu+L}$linear
Go deeper: can any gradient method beat $\frac{Q_f-1}{Q_f+1}$?

Yes. [Y] Theorem 2.1.13 shows that no first-order method can beat $\big(\frac{\sqrt{Q_f}-1}{\sqrt{Q_f}+1}\big)^{2k}$ in the worst case, and that order is achievable. Replacing $Q_f$ by $\sqrt{Q_f}$ is a huge saving when $Q_f$ is large. On quadratics, the conjugate gradient method (Parts 8–9) achieves exactly this $\sqrt\kappa$ behaviour; Newton-type methods (Part 10) remove the dependence on $\kappa$ near the solution.

$f\in\mathcal S_{\mu,L}^{1,1}$ with $\mu=1$, $L=9$. Give the step $h=2/(\mu+L)$ and the distance contraction factor $\frac{Q_f-1}{Q_f+1}$.

$Q_f=L/\mu=9$.

$h=2/10=0.2$ and $\frac{9-1}{9+1}=0.8$.

Same $f$ ($\mu=1$, $L=9$), $h=0.2$, $\norm{\x_0-\x^\star}=10$. What is the smallest $k$ for which the theorem guarantees $\norm{\x_k-\x^\star}\le10^{-3}$?

Solve $10\cdot0.8^k\le10^{-3}$, i.e. $0.8^k\le10^{-4}$.

$k\ge\ln10^4/\ln1.25=9.2103/0.22314=41.28$, so $k=42$. (Check: $10\cdot0.8^{41}\approx1.06\times10^{-3}$, $10\cdot0.8^{42}\approx0.85\times10^{-3}$.)

Same $f$ ($\mu=1$, $L=9$) but with the step $h=1/L$. What per-step factor does Theorem 2.1.15 guarantee for the distance $\norm{\x_k-\x^\star}$ (not its square)?

First check $h\le2/(\mu+L)$. Then the factor on $r_k^2$ is $1-\frac{2h\mu L}{\mu+L}$; take its square root.

$1/9\le0.2$, so the theorem applies. $1-\frac{2\cdot\frac19\cdot1\cdot9}{10}=1-0.2=0.8$ on $r_k^2$, so $\sqrt{0.8}\approx0.894$ on $r_k$, slower than the $0.8$ of the optimal step. In general $h=1/L$ gives $\sqrt{\frac{Q_f-1}{Q_f+1}}$.

Same $f$ ($\mu=1$, $L=9$, $h=0.2$, $\norm{\x_0-\x^\star}=10$). What bound does the theorem give on $f(\x_{10})-f^\star$?

$\frac L2\big(\frac{Q_f-1}{Q_f+1}\big)^{2k}r_0^2$ with $k=10$.

$4.5\cdot0.8^{20}\cdot100=4.5\cdot0.011529\cdot100\approx5.188$.

  • Read off $\mu$ and $L$ from Hessian bounds $\mu I\preceq\hess f\preceq LI$
  • Use $h=2/(\mu+L)$ for the best guaranteed factor $\frac{Q_f-1}{Q_f+1}$
  • Square the factor for function values: $\big(\frac{Q_f-1}{Q_f+1}\big)^{2k}$, with the constant $\frac L2r_0^2$
  • Estimate work as $\approx\frac{Q_f}2\ln\frac{r_0}\varepsilon$ iterations
  • Applying Theorem 2.1.15 with $h>2/(\mu+L)$ (the theorem says nothing there)
  • Quoting a linear rate for a function that is only convex ($\mu=0$)
  • Forgetting the square root when converting the bound on $r_k^2$ into one on $r_k$
  1. Strong co-coercivity: $(\grad f(\x)-\grad f(\y))^\top(\x-\y)\ge\frac{\mu L}{\mu+L}\norm{\x-\y}^2+\frac1{\mu+L}\norm{\grad f(\x)-\grad f(\y)}^2$.
  2. For $f\in\mathcal S_{\mu,L}^{1,1}$ and $h=\frac2{\mu+L}$: $\norm{\x_k-\x^\star}\le\big(\frac{Q_f-1}{Q_f+1}\big)^k r_0$ and $f(\x_k)-f^\star\le\frac L2\big(\frac{Q_f-1}{Q_f+1}\big)^{2k}r_0^2$.
  3. Strong convexity upgrades $O(1/k)$ to a linear rate, with the same factor Part 5 found for quadratics; the work grows in proportion to $Q_f$.

For $f(x,y)=x^2+5y^2$, gradient descent with $h=2/(\mu+L)$ contracts the distance to the minimizer per step by at least…

$4/5$
That would be $(L-\mu)/L$. Find $\mu$ and $L$ from the Hessian $\mathrm{diag}(2,10)$ and use $\frac{Q_f-1}{Q_f+1}$.
$2/3$
$\mu=2$, $L=10$, $Q_f=5$, $\frac{5-1}{5+1}=\frac23$.
$4/9$
That is $(2/3)^2$, the factor for function values, not distances.

What makes the proof of Theorem 2.1.15 give a fixed-fraction contraction, unlike Theorem 2.1.14?

A smaller step size
Both theorems use comparable steps; the difference is in the inequality used.
The term $\frac{\mu L}{\mu+L}\norm{\x_k-\x^\star}^2$ in strong co-coercivity
It subtracts a multiple of $r_k^2$ itself, giving $r_{k+1}^2\le(1-c)r_k^2$.
Using exact line search
The theorem uses a constant step; no line search is involved.

If $Q_f$ grows from 10 to 1000, the number of iterations Theorem 2.1.15 needs for a fixed accuracy grows by roughly…

a factor of 10, because of a square root
A square root appears for accelerated or conjugate gradient methods, not for plain gradient descent.
it stays about the same
$\ln(1/\rho)\approx2/Q_f$ depends on $Q_f$.
a factor of about 100
Iterations $\approx\frac{Q_f}2\ln\frac{r_0}\varepsilon$, proportional to $Q_f$.

Theorem 2.1.15's guaranteed factor on $\norm{\x_k-\x^\star}$, as a function of $h\in(0,2/(\mu+L)]$, is…

decreasing in $h$, best at $h=2/(\mu+L)$
$\sqrt{1-\frac{2h\mu L}{\mu+L}}$ decreases as $h$ grows.
smallest at $h=1/L$
$1/L\le2/(\mu+L)$, and the factor keeps decreasing up to $2/(\mu+L)$.
independent of $h$
Look at $1-\frac{2h\mu L}{\mu+L}$.

Method A has $f(\x_k)-f^\star\le C/k$; method B has $\norm{\x_k-\x^\star}\le0.999^kr_0$. Which statement is right?

A is faster at every accuracy
For very small $\varepsilon$, $C/\varepsilon$ eventually dwarfs $\ln(1/\varepsilon)/\ln(1/0.999)$.
B is linear and wins for small enough $\varepsilon$, but A can win at moderate accuracy
B needs about $1000\ln(r_0/\varepsilon)$ steps, A about $C/\varepsilon$; which is smaller depends on $\varepsilon$ and $C$.
Both are linear
$C/k$ has $e_{k+1}/e_k\to1$: sublinear.

The inequality $f(\x_{k+1})\le f(\x_k)-h(1-\frac12hL)\norm{\g_k}^2$ for $\x_{k+1}=\x_k-h\g_k$ comes from…

convexity
No convexity is needed; it holds for non-convex $f$ too.
the Descent Lemma $f(\y)\le f(\x)+\grad f(\x)^\top(\y-\x)+\frac L2\norm{\y-\x}^2$
Put $\y-\x=-h\g_k$.
the Wolfe curvature condition
There's no line search here: the step is constant.

Which function belongs to $\mathcal F_L^{1,1}$ for some $L$ but to no $\mathcal S_{\mu,L}^{1,1}$ with $\mu>0$?

$f(x)=x^2$
$f''=2$: strongly convex with $\mu=2$.
$f(x)=x^4$
$f''=12x^2$ is unbounded, so $f'$ is not globally Lipschitz.
$f(x)=\ln(1+e^x)$
$f''=\sigma(1-\sigma)\in(0,\frac14]$: convex and $\frac14$-smooth, but $f''\to0$, so no $\mu>0$ works. (It also has no minimizer.)

A method uses $\d_k=-B_k^{-1}\g_k$ with $B_k\succ0$ but $\kappa(B_k)=k+1$ growing without bound. What does Zoutendijk's theory guarantee?

$\norm{\g_k}\to0$, since every $\d_k$ is a descent direction
Descent alone isn't enough: the bound $\cos\theta_k\ge1/\kappa(B_k)$ now tends to 0.
Only $\sum_k\cos^2\theta_k\norm{\g_k}^2\lt\infty$; the corollary doesn't apply
There's no fixed $\delta>0$ with $\cos\theta_k\ge\delta$, so $\norm{\g_k}\to0$ is not guaranteed.
Divergence of $\x_k$
The theorem says nothing of the kind; it just loses its conclusion.

For the fit-the-line loss ($\kappa\approx46.2$), gradient descent with $h=2/(m+L)$ shrinks the distance to $(1.5,1/3)$ per step by about…

$0.5$
Compute $\frac{\kappa-1}{\kappa+1}$ with $\kappa=46.2$.
$0.979$
That's $1-1/\kappa$; the optimal-step rate is $\frac{\kappa-1}{\kappa+1}$.
$0.958$
$\frac{45.2}{47.2}\approx0.958$: about $\frac\kappa2\ln10\approx53$ steps per extra digit.

Lectures 11 and 12 (17 and 22 September): how to beat gradient descent's zig-zag on a quadratic by choosing search directions that never undo each other. You will meet conjugacy (orthogonality in the geometry of $A$), prove that $n$ such directions solve an $n$-variable quadratic in exactly $n$ exact line searches, prove the Expanding Subspace Theorem, and then build and prove correct the conjugate gradient (CG) method, one of the most important algorithms in scientific computing.

You need: Part 0b (eigenvalues, positive definite matrices, the spectral theorem), Part 4 (a stationary point of a convex function is a global minimizer) and Part 5 (gradient descent on quadratics, exact line search, zig-zag, the condition number $\kappa$). Part 9 continues with how fast CG converges when $n$ is huge.

Conjugacy: directions that are orthogonal after stretching

Two directions $\d_i,\d_j$ are conjugate with respect to $A$ when $\d_i^\top A\d_j=0$: they are perpendicular once you stretch space so that the elliptical contours of the quadratic become circles.

Gradient descent zig-zags because each step partly undoes the previous one. Conjugate directions are exactly the directions that never undo each other, and they are the foundation of both theorems in this part.

Tidying a room one shelf at a time works only if putting books on shelf 2 never knocks books off shelf 1. Conjugate directions are shelves that don't interfere.

The problem for this whole part

Throughout, $A$ is a symmetric positive definite $n\times n$ matrix ($A=A^\top\succ0$, Part 0b) and $$f(\x)=\tfrac12\x^\top A\x-\b^\top\x,\qquad \g(\x)=\grad f(\x)=A\x-\b.$$ Because $A\succ0$, $f$ is strictly convex with the unique minimizer $\x^\star=A^{-1}\b$ (the point where $\g=\0$). So minimizing $f$ and solving the linear system $A\x=\b$ are the same problem. We write $\x_k$ for the iterates and $\g_k=\g(\x_k)=A\x_k-\b$.

Careful with signs and letters across books. [NW] calls $A\x_k-\b$ the residual $\r_k$ and uses $\p_k$ for directions, so [NW]'s $\r_k$ is our $\g_k$. [LD] writes $\g_k$ (as we do) and $Q$ for $A$. Many numerical linear algebra texts instead define the residual as $\b-A\x_k=-\g_k$, the opposite sign. If a formula from another source has the "wrong" sign, check which convention it uses before anything else.

Why something better than gradient descent?

Recall from Part 5: gradient descent with exact line search on this $f$ makes consecutive gradients orthogonal ($\g_{k+1}^\top\g_k=0$), so on a stretched bowl it zig-zags, and the error shrinks only by a factor tied to $\kappa=L/m$ each step. Even on a $2\times2$ problem it generally never lands exactly on $\x^\star$.

Here is a hint of what could work. If $A$ were diagonal, $f$ would be a sum of separate one-variable parabolas, $f=\sum_i(\tfrac12a_{ii}x_i^2-b_ix_i)$. Minimizing exactly along $x_1$, then $x_2$, …, then $x_n$ would finish in $n$ steps, because fixing $x_2$ can't spoil the choice of $x_1$. The contours are ellipses whose axes line up with the coordinate axes. For a general $A$ the ellipses are tilted, and line searches along the coordinate axes interfere. The fix: find $n$ directions that behave, for this $A$, the way the coordinate axes behave for a diagonal matrix.

Take $A=\begin{pmatrix}4&1\\1&3\end{pmatrix}$, $\b=(1,2)^\top$ (so $\x^\star=(\tfrac1{11},\tfrac7{11})^\top$), start at $\x_0=(2,1)^\top$, and do an exact line search along $\e_1$, then one along $\e_2$. Where do you end up?

Exactly at $\x^\star$, because two perpendicular directions span $\R^2$
Spanning $\R^2$ is not enough: the second search can spoil the first.
Near $\x^\star$ but not on it
The first search gives $\x_1=(0,1)$; the second gives $\x_2=(0,\tfrac23)$, while $\x^\star\approx(0.091,0.636)$. The second step changed the slope along $\e_1$.
Further from $\x^\star$ than $\x_0$
Each exact line search lowers $f$, so you can't do worse than where you started (in $f$).
Try it

This is the coordinate-axes experiment. Press Step twice and look at $\g_2^\top\d_0$: it is not zero, so after step 2 you could still go downhill along $\e_1$. Then switch to "your $\d_0$ + its conjugate" and step again: whatever angle you pick for $\d_0$, the pair lands exactly on the gold star.

Let $A\succ0$. Nonzero vectors $\d_0,\dots,\d_{m-1}\in\R^n$ are $A$-conjugate (or $A$-orthogonal) if $$\d_i^\top A\d_j=0\qquad\text{for all }i\ne j.$$ Equivalently, they are orthogonal in the $A$-inner product $\ip{\u}{\v}_A=\u^\top A\v$, whose norm is $\norm{\u}_A=\sqrt{\u^\top A\u}$. This is a genuine inner product because $A$ is symmetric and $\u^\top A\u\gt0$ for $\u\ne\0$.

With $A=I$, conjugate just means orthogonal. For any other $A$ the two ideas differ. For $A=\begin{pmatrix}4&1\\1&3\end{pmatrix}$, the vectors $\d_0=(1,0)^\top$ and $\d_1=(-\tfrac14,1)^\top$ are conjugate: $A\d_1=(-1+1,\,-\tfrac14+3)^\top=(0,\tfrac{11}4)^\top$, so $\d_0^\top A\d_1=0$. But $\d_0^\top\d_1=-\tfrac14\ne0$: they are not perpendicular.

The picture: stretch the ellipses into circles

Write $A=V\Lambda V^\top$ (spectral theorem) and let $A^{1/2}=V\Lambda^{1/2}V^\top$, the symmetric positive definite matrix with $A^{1/2}A^{1/2}=A$. Change variables to $\y=A^{1/2}\x$. Then $$\u^\top A\v=(A^{1/2}\u)^\top(A^{1/2}\v),\qquad \x^\top A\x=\norm{\y}^2.$$ So the level sets $\x^\top A\x=c$ (ellipses) become circles $\norm{\y}^2=c$, and $\u,\v$ are $A$-conjugate exactly when their stretched images $A^{1/2}\u$, $A^{1/2}\v$ are perpendicular. Conjugacy is ordinary orthogonality, seen through the lens that makes the bowl round.

A second picture you can draw by hand: take the ellipse $\x^\top A\x=c$ through the tip of $\d_0$. The gradient of $\x^\top A\x$ at that tip is $2A\d_0$, which is perpendicular to the tangent. So the tangent direction $\v$ satisfies $\v^\top A\d_0=0$: the direction conjugate to $\d_0$ is parallel to the tangent of the ellipse where $\d_0$ meets it. In $\R^2$ that pins it down up to scaling; in $\R^n$ the vectors conjugate to $\d_0$ form an $(n-1)$-dimensional subspace, $\{\v:\v^\top(A\d_0)=0\}$.

Try it

Drag the tip of $\d_0$ (blue). The purple arrow is the conjugate direction $\d_1$, and the gold dashed line is the tangent to the blue ellipse at the tip of $\d_0$: they are always parallel. Read $\cos\angle(\d_0,\d_1)$: usually not 0. Now press "Animate the stretch": as the ellipses become circles, the on-screen angle becomes exactly 90°. Finally, snap $\d_0$ to an eigenvector: then $\d_0$ and $\d_1$ are both orthogonal and conjugate. On the diagonal matrix, check that $\e_1$ and $\e_2$ are conjugate.

If $\d_0,\dots,\d_{m-1}$ are nonzero and $A$-conjugate with $A\succ0$, they are linearly independent. In particular $m\le n$: there are at most $n$ nonzero mutually conjugate vectors in $\R^n$.

Proof (examinable). Suppose $\sum_{i=0}^{m-1}c_i\d_i=\0$. Multiply on the left by $\d_j^\top A$. Every term with $i\ne j$ vanishes by conjugacy, leaving $c_j\,\d_j^\top A\d_j=0$. Since $\d_j\ne\0$ and $A\succ0$, $\d_j^\top A\d_j\gt0$, so $c_j=0$. This holds for every $j$, so the only combination giving $\0$ is the trivial one. $\blacksquare$

The proof is the "multiply by $\d_j^\top A$ to kill every term but one" trick. You'll use it again three times in this part.

Expanding $\x^\star$ in a conjugate basis

Let $\d_0,\dots,\d_{n-1}$ be $A$-conjugate. By the lemma they form a basis of $\R^n$, so $\x^\star=\sum_ia_i\d_i$ for some numbers $a_i$. Multiply by $\d_j^\top A$: $$\d_j^\top A\x^\star=a_j\,\d_j^\top A\d_j\quad\Longrightarrow\quad a_j=\frac{\d_j^\top A\x^\star}{\d_j^\top A\d_j}=\frac{\d_j^\top\b}{\d_j^\top A\d_j},$$ using $A\x^\star=\b$. The last expression needs only $\b$, not the unknown $\x^\star$. That is the whole point of $A$-orthogonality. With an ordinary orthogonal basis the coefficient would be $\d_j^\top\x^\star/\norm{\d_j}^2$, which needs $\x^\star$, the thing we're trying to find. ([LD] §9.1 p.264.)

Where do conjugate directions come from?

  • Eigenvectors. The orthonormal eigenvectors $\v_1,\dots,\v_n$ of $A$ satisfy $\v_i^\top A\v_j=\lambda_j\v_i^\top\v_j=0$ for $i\ne j$: both orthogonal and conjugate. But computing them is harder than solving $A\x=\b$.
  • Gram–Schmidt in the $A$-inner product. From any basis $\u_0,\dots,\u_{n-1}$: $\d_0=\u_0$ and $$\d_k=\u_k-\sum_{i\lt k}\frac{\u_k^\top A\d_i}{\d_i^\top A\d_i}\,\d_i .$$ Each correction removes the $A$-component along one old direction. This works, but it stores all old directions and costs $O(n^3)$ for dense $A$. The conjugate gradient method of Chapter 8.3 is a way to do it almost for free.

Let $A=\begin{pmatrix}2&1\\1&3\end{pmatrix}$ and $\b=(3,4)^\top$. (a) Find a direction $\d_1=(1,c)^\top$ conjugate to $\d_0=(1,0)^\top$. (b) Write $\x^\star=a_0\d_0+a_1\d_1$ without solving $A\x=\b$ first, then check.

  1. $\d_0^\top A\d_1=(1,0)\begin{pmatrix}2+c\\1+3c\end{pmatrix}=2+c$. Setting it to 0 gives $c=-2$, so $\d_1=(1,-2)^\top$.

    Conjugacy to $\d_0$ is one linear equation $\v^\top(A\d_0)=0$ with $A\d_0=(2,1)^\top$; in $\R^2$ it fixes the direction.

  2. $\d_0^\top A\d_0=2$ and $\d_1^\top A\d_1=(1,-2)\,(0,-5)^\top=10$. Also $\d_0^\top\b=3$ and $\d_1^\top\b=3-8=-5$.

    These are the only ingredients of $a_j=\d_j^\top\b/\d_j^\top A\d_j$.

  3. $a_0=\tfrac32$, $a_1=-\tfrac5{10}=-\tfrac12$, so $\x^\star=\tfrac32(1,0)^\top-\tfrac12(1,-2)^\top=(1,1)^\top$.

    We never inverted $A$: conjugacy decoupled the coefficients.

  4. Check: $A(1,1)^\top=(3,4)^\top=\b$. ✓ Note $\d_0^\top\d_1=1\ne0$: conjugate but not orthogonal.

    Always verify with one matrix–vector product; it catches sign slips.

For $A=\begin{pmatrix}5&2\\2&1\end{pmatrix}$, find $c$ so that $(c,1)^\top$ is $A$-conjugate to $(1,1)^\top$.

Compute $A(1,1)^\top$ first, then require $(c,1)$ to be orthogonal to it.

$A(1,1)^\top=(5+2,\,2+1)^\top=(7,3)^\top$. Then $(c,1)\cdot(7,3)=7c+3=0$, so $c=-\tfrac37\approx-0.4286$.

Which pairs are $A$-conjugate? (i) $(1,1)^\top,(1,-1)^\top$ for $A=\begin{pmatrix}3&1\\1&3\end{pmatrix}$. (ii) $(1,1)^\top,(1,-1)^\top$ for $A=\begin{pmatrix}2&1\\1&3\end{pmatrix}$.

Compute $A(1,-1)^\top$ in each case and dot it with $(1,1)$.

(i) $A(1,-1)^\top=(2,-2)^\top$, and $(1,1)\cdot(2,-2)=0$: conjugate (they are the eigenvectors of this $A$). (ii) $A(1,-1)^\top=(1,-2)^\top$, and $(1,1)\cdot(1,-2)=-1\ne0$: not conjugate, even though the two vectors are perpendicular.

Apply Gram–Schmidt in the $A$-inner product to $\e_1,\e_2,\e_3$ with $A=\begin{pmatrix}2&-1&0\\-1&2&-1\\0&-1&2\end{pmatrix}$. You get $\d_0=\e_1$ and $\d_1=(\tfrac12,1,0)^\top$. Find $\d_2$.

$\d_2=\e_3-\frac{\e_3^\top A\d_0}{\d_0^\top A\d_0}\d_0-\frac{\e_3^\top A\d_1}{\d_1^\top A\d_1}\d_1$. Compute $A\d_1$ first.

$A\d_0=(2,-1,0)^\top$, so $\e_3^\top A\d_0=0$: no correction along $\d_0$. $A\d_1=(0,\tfrac32,-1)^\top$, so $\e_3^\top A\d_1=-1$ and $\d_1^\top A\d_1=\tfrac32$. Hence $\d_2=\e_3+\tfrac23\d_1=(\tfrac13,\tfrac23,1)^\top$. Check: $\d_2^\top A\d_0=\tfrac23-\tfrac23=0$ and $\d_2^\top A\d_1=\tfrac23\cdot\tfrac32-1=0$.

What is the largest number of nonzero, mutually $A$-conjugate vectors that can exist in $\R^4$ (for a given $A\succ0$)?

What does the lemma say about conjugate vectors and linear independence?

Nonzero conjugate vectors are linearly independent, and $\R^4$ holds at most 4 linearly independent vectors. Four are achievable (for example the eigenvectors). So the answer is 4.

  • Test conjugacy by computing $A\d_j$ once and dotting it with $\d_i$
  • Use $a_j=\d_j^\top\b/\d_j^\top A\d_j$ to expand $\x^\star$ without knowing it
  • Picture conjugate directions as perpendicular after stretching by $A^{1/2}$
  • State which sign convention ($\g=A\x-\b$ or $\r=\b-A\x$) you are using
  • Confusing orthogonal ($\d_i^\top\d_j=0$) with conjugate ($\d_i^\top A\d_j=0$)
  • Forgetting that conjugacy needs $A\succ0$ (otherwise $\d^\top A\d$ can be 0 for $\d\ne\0$ and independence fails)
  • Thinking eigenvectors are a practical source of conjugate directions for large problems
  1. $\d_i,\d_j$ are $A$-conjugate when $\d_i^\top A\d_j=0$: orthogonal in the $A$-inner product, i.e. perpendicular after the change of variables $\y=A^{1/2}\x$.
  2. Nonzero conjugate vectors are linearly independent (multiply by $\d_j^\top A$), so at most $n$ exist and $n$ of them form a basis.
  3. In a conjugate basis, $\x^\star=\sum_j\frac{\d_j^\top\b}{\d_j^\top A\d_j}\d_j$: every coefficient is computable without knowing $\x^\star$.

For $A\succ0$ that is not a multiple of $I$, which statement is true?

Orthogonal vectors are always $A$-conjugate
Try $(1,1)$ and $(1,-1)$ with $A=\begin{pmatrix}2&1\\1&3\end{pmatrix}$.
$A$-conjugate vectors are always orthogonal
Check $(1,0)$ and $(1,-2)$ for $A=\begin{pmatrix}2&1\\1&3\end{pmatrix}$.
The eigenvectors of $A$ are both orthogonal and $A$-conjugate
$\v_i^\top A\v_j=\lambda_j\v_i^\top\v_j=0$ for $i\ne j$.

Why does the expansion $\x^\star=\sum a_j\d_j$ use $A$-conjugate rather than orthogonal $\d_j$?

Because orthogonal vectors don't form a basis
$n$ nonzero orthogonal vectors do form a basis. The issue is about computing the coefficients.
Because then $a_j=\d_j^\top\b/\d_j^\top A\d_j$ involves only known data, not $\x^\star$
Multiplying by $\d_j^\top A$ turns $\d_j^\top A\x^\star$ into $\d_j^\top\b$.
Because conjugate vectors have unit $A$-norm
Conjugacy says nothing about lengths; the formula divides by $\d_j^\top A\d_j$ anyway.

In the proof that conjugate vectors are independent, where is $A\succ0$ used?

To make $A$ symmetric
Symmetry is a separate assumption. Look at the very last step of the proof.
To kill the cross terms $\d_j^\top A\d_i$
Those vanish by conjugacy, not by positive definiteness.
To conclude $c_j=0$ from $c_j\,\d_j^\top A\d_j=0$, since $\d_j^\top A\d_j\gt0$
If $\d_j^\top A\d_j$ could be 0, you couldn't divide by it.

In the stretched coordinates $\y=A^{1/2}\x$, the function $\tfrac12\x^\top A\x$ becomes…

$\tfrac12\norm{\y}^2$, whose level sets are circles
$\x^\top A\x=(A^{1/2}\x)^\top(A^{1/2}\x)=\norm{\y}^2$.
$\tfrac12\y^\top A^2\y$
Substitute $\x=A^{-1/2}\y$ and simplify $A^{-1/2}AA^{-1/2}$.
$\tfrac12\y^\top A^{-1}\y$
The stretch is designed to remove $A$ entirely. Recompute.

The conjugate direction method and the Expanding Subspace Theorem

Do one exact line search along each of $n$ conjugate directions, in any order, and you land exactly on $\x^\star$; along the way, each $\x_k$ is the best point in the whole subspace explored so far.

These two theorems ([NW] Thm 5.1 and 5.2) are the backbone of CG. The Expanding Subspace Theorem's three-line induction is the most reused proof in this block of the course and a standard exam question.

Exploring a building floor by floor: once you've found the best spot on floors 1 and 2, adding floor 3 means you only need to search the new floor, never re-search the old ones.

Given $A$-conjugate directions $\d_0,\dots,\d_{n-1}$ and any $\x_0$, for $k=0,1,\dots,n-1$: $$\g_k=A\x_k-\b,\qquad \alpha_k=-\frac{\g_k^\top\d_k}{\d_k^\top A\d_k},\qquad \x_{k+1}=\x_k+\alpha_k\d_k.$$

$\alpha_k$ is the unique minimizer of $\phi(\alpha)=f(\x_k+\alpha\d_k)$. Moreover $$\g_{k+1}=\g_k+\alpha_kA\d_k\qquad\text{and}\qquad\g_{k+1}^\top\d_k=0.$$

Proof (examinable). Expanding, $\phi(\alpha)=f(\x_k)+\alpha\,\g_k^\top\d_k+\tfrac{\alpha^2}2\d_k^\top A\d_k$, a parabola in $\alpha$ with positive leading coefficient. $\phi'(\alpha)=\g_k^\top\d_k+\alpha\,\d_k^\top A\d_k=0$ gives $\alpha_k$. Next, $\g_{k+1}=A(\x_k+\alpha_k\d_k)-\b=\g_k+\alpha_kA\d_k$. Finally, by the chain rule $\phi'(\alpha_k)=\g_{k+1}^\top\d_k$, and this is 0 at the minimizer. $\blacksquare$

In words: after an exact line search, the new gradient is perpendicular to the direction you just searched. (You met this in Part 5 for $\d_k=-\g_k$.) Note that $\alpha_k$ may be negative: the method walks along the line, in whichever sense goes down.

For any $\x_0\in\R^n$, the conjugate direction method reaches the minimizer in at most $n$ steps: $\x_n=\x^\star$.

Prove the Conjugate Direction Theorem.

  1. The $\d_i$ form a basis, so $\x^\star-\x_0=\sum_{i=0}^{n-1}\sigma_i\d_i$ for some numbers $\sigma_i$, and multiplying by $\d_k^\top A$ gives $\sigma_k=\dfrac{\d_k^\top A(\x^\star-\x_0)}{\d_k^\top A\d_k}$.

    The independence lemma and the "multiply by $\d_k^\top A$" trick from Chapter 8.1.

  2. Unrolling the iteration, $\x_k-\x_0=\sum_{i\lt k}\alpha_i\d_i$, so by conjugacy $\d_k^\top A(\x_k-\x_0)=0$.

    Every term is $\alpha_i\d_k^\top A\d_i$ with $i\ne k$.

  3. Therefore $\d_k^\top A(\x^\star-\x_0)=\d_k^\top A(\x^\star-\x_k)+\d_k^\top A(\x_k-\x_0)=\d_k^\top(\b-A\x_k)=-\d_k^\top\g_k$.

    Insert $\pm\x_k$, use step 2 for the second piece and $A\x^\star=\b$ for the first.

  4. So $\sigma_k=-\g_k^\top\d_k/\d_k^\top A\d_k=\alpha_k$ for every $k$, and $\x_n=\x_0+\sum_{i\lt n}\alpha_i\d_i=\x_0+\sum_{i\lt n}\sigma_i\d_i=\x^\star$. $\blacksquare$

    Each step adds exactly the right amount of its own direction, and later steps never change that amount.

Why it works: coordinate descent in disguise

Put the directions in the columns of $S=[\d_0|\cdots|\d_{n-1}]$ and write $\x=S\hat\x$. Then $f(S\hat\x)=\tfrac12\hat\x^\top(S^\top AS)\hat\x-(S^\top\b)^\top\hat\x$, and $S^\top AS=\mathrm{diag}(\d_0^\top A\d_0,\dots,\d_{n-1}^\top A\d_{n-1})$ by conjugacy. In the $\hat\x$ coordinates the function is a sum of $n$ separate parabolas, and step $k$ minimizes exactly along the $k$-th coordinate axis. That's the diagonal case from Chapter 8.1: no later move can disturb an earlier one. When $A$ is already diagonal, the coordinate axes themselves are conjugate.

What is true at every intermediate step

Let $\mathcal B_k=\operatorname{span}\{\d_0,\dots,\d_{k-1}\}$ (with $\mathcal B_0=\{\0\}$). These subspaces grow by one dimension per step, $\mathcal B_1\subset\mathcal B_2\subset\cdots\subset\mathcal B_n=\R^n$, and $\x_k$ lies in the affine set $\x_0+\mathcal B_k$ ("$\x_0$ plus anything in $\mathcal B_k$"; [LD] calls it a linear variety).

Let $\d_0,\dots,\d_{n-1}$ be $A$-conjugate, $\x_0$ arbitrary, and $\x_k$ generated by the conjugate direction method. Then for each $k=1,\dots,n$:

  1. $\g_k^\top\d_i=0$ for all $i=0,1,\dots,k-1$, i.e. $\g_k$ is orthogonal to $\mathcal B_k$;
  2. $\x_k$ is the unique minimizer of $f$ over the affine set $\x_0+\mathcal B_k$ (and also over the line $\{\x_{k-1}+\alpha\d_{k-1}\}$).

Prove part (i) of the Expanding Subspace Theorem by induction on $k$.

  1. Base case $k=1$: $\g_1^\top\d_0=0$ by the exact-line-search lemma.

    The first step is an exact line search along $\d_0$, so the new gradient is perpendicular to $\d_0$.

  2. Inductive step. Assume $\g_k^\top\d_i=0$ for all $i\lt k$. From the lemma, $\g_{k+1}=\g_k+\alpha_kA\d_k$, so for any $i\le k$: $$\g_{k+1}^\top\d_i=\g_k^\top\d_i+\alpha_k\,\d_k^\top A\d_i.$$

    The gradient update is the engine of the proof: the new gradient is the old one plus a multiple of $A\d_k$ (using $A^\top=A$ to write $(A\d_k)^\top\d_i=\d_k^\top A\d_i$).

  3. Case $i=k$: $\g_{k+1}^\top\d_k=0$ directly by the lemma (exact line search along $\d_k$).

    The newest direction is handled by the line search itself.

  4. Case $i\lt k$: $\g_k^\top\d_i=0$ by the induction hypothesis, and $\d_k^\top A\d_i=0$ by conjugacy. So $\g_{k+1}^\top\d_i=0$.

    The older directions are protected by conjugacy: moving along $\d_k$ changes the gradient by $\alpha_kA\d_k$, which has no component along any old $\d_i$ in the dot product. This is exactly where orthogonal-but-not-conjugate directions fail.

  5. So $\g_{k+1}\perp\d_0,\dots,\d_k$, i.e. $\g_{k+1}\perp\mathcal B_{k+1}$, completing the induction. $\blacksquare$

    Write the conclusion in the subspace language, because part (ii) uses it in that form.

Proof of (ii). Write points of $\x_0+\mathcal B_k$ as $\x_0+D_k\y$ with $D_k=[\d_0|\cdots|\d_{k-1}]$ and $\y\in\R^k$, and let $h(\y)=f(\x_0+D_k\y)$. This is a quadratic in $\y$ with Hessian $D_k^\top AD_k=\mathrm{diag}(\d_i^\top A\d_i)\succ0$, so it is strictly convex. By the chain rule $\grad h(\y)=D_k^\top\grad f(\x_0+D_k\y)$. The iterate $\x_k=\x_0+\sum_{i\lt k}\alpha_i\d_i$ corresponds to $\y=(\alpha_0,\dots,\alpha_{k-1})$, where $\grad h=D_k^\top\g_k=\0$ by part (i). A stationary point of a convex function is a global minimizer (Part 4), and strict convexity makes it unique. The line through $\x_{k-1}$ along $\d_{k-1}$ lies inside $\x_0+\mathcal B_k$ and contains $\x_k$, so $\x_k$ minimizes over that line too. $\blacksquare$

Taking $k=n$: $\x_n$ minimizes $f$ over $\x_0+\R^n=\R^n$, so $\x_n=\x^\star$. That's a second proof of the Conjugate Direction Theorem.

Another reading: the closest point in the $A$-norm

Using $A\x^\star=\b$, a short expansion gives $f(\x)-f(\x^\star)=\tfrac12\norm{\x-\x^\star}_A^2$. So minimizing $f$ is the same as minimizing the $A$-norm distance to the solution, and the theorem says: $\x_k$ is the point of $\x_0+\mathcal B_k$ closest to $\x^\star$ in the $A$-norm. Since $\g_k=A(\x_k-\x^\star)$, part (i) reads $\ip{\x_k-\x^\star}{\d_i}_A=0$: the error is $A$-orthogonal to the explored subspace, just like "the best approximation from a subspace is the orthogonal projection". Part 9 builds its convergence analysis on this view.

Go deeper: checking $f(\x)-f^\star=\tfrac12\norm{\x-\x^\star}_A^2$

$\tfrac12(\x-\x^\star)^\top A(\x-\x^\star)=\tfrac12\x^\top A\x-\x^\top A\x^\star+\tfrac12\x^{\star\top}A\x^\star=\tfrac12\x^\top A\x-\b^\top\x+\tfrac12\b^\top\x^\star$. And $f^\star=f(\x^\star)=\tfrac12\b^\top\x^\star-\b^\top\x^\star=-\tfrac12\b^\top\x^\star$. So the right side is $f(\x)-f^\star$. ([LD] §9.2, Fig. 9.2.)

Run the conjugate direction method on $A=\begin{pmatrix}4&1\\1&3\end{pmatrix}$, $\b=(1,2)^\top$ with $\d_0=(1,0)^\top$, $\d_1=(-\tfrac14,1)^\top$, from $\x_0=(2,1)^\top$.

  1. $\g_0=A\x_0-\b=(9-1,\,5-2)^\top=(8,3)^\top$, $\d_0^\top A\d_0=4$, so $\alpha_0=-\g_0^\top\d_0/4=-2$ and $\x_1=(0,1)^\top$.

    Negative $\alpha$ just means we move left along $\d_0$; that's downhill here since $\g_0^\top\d_0=8\gt0$.

  2. $\g_1=A\x_1-\b=(1,3)^\top-(1,2)^\top=(0,1)^\top$. Check $\g_1^\top\d_0=0$. ✓

    Expanding Subspace (i) at $k=1$.

  3. $A\d_1=(0,\tfrac{11}4)^\top$, so $\d_1^\top A\d_1=\tfrac{11}4$; $\g_1^\top\d_1=1$; $\alpha_1=-\tfrac4{11}$.

    Same formula, next direction.

  4. $\x_2=(0,1)^\top-\tfrac4{11}(-\tfrac14,1)^\top=(\tfrac1{11},\tfrac7{11})^\top=\x^\star$, and $\g_2=\0$. Done in $n=2$ steps.

    Compare the coordinate axes from the same start, which end at $(0,\tfrac23)$ (the predict question above).

Try it

A conjugate direction run on a $3\times3$ problem, with directions from $A$-Gram–Schmidt applied to a basis you choose. Step through it and watch the checks: the green zeros $\g_k^\top\d_i$ ($i\lt k$) are part (i); the matching values of $f(\x_k)$ and the directly computed minimum over $\x_0+\mathcal B_k$ are part (ii). The two bases give different directions and different intermediate points, but both end at $\x^\star$ after 3 steps.

Same $A=\begin{pmatrix}4&1\\1&3\end{pmatrix}$, $\b=(1,2)^\top$, $\x_0=(2,1)^\top$, but use the orthogonal directions $\e_1$ then $\e_2$ with exact line searches. Find $\x_2$.

Step 1 is identical to the worked example ($\x_1=(0,1)$, $\g_1=(0,1)$). For step 2, $\alpha=-\g_1^\top\e_2/\e_2^\top A\e_2$.

$\alpha_1=-1/3$, so $\x_2=(0,\tfrac23)^\top\ne\x^\star$. Indeed $\g_2=A\x_2-\b=(\tfrac23-1,\,2-2)^\top=(-\tfrac13,0)^\top$, and $\g_2^\top\e_1=-\tfrac13\ne0$: the induction breaks at the term $\alpha_1\e_2^\top A\e_1=-\tfrac13\cdot1$, because $\e_1^\top A\e_2=1\ne0$.

Run the conjugate direction method on $A=\begin{pmatrix}2&1\\1&3\end{pmatrix}$, $\b=(3,4)^\top$, from $\x_0=\0$ with $\d_0=(1,0)^\top$, $\d_1=(1,-2)^\top$. Find $\alpha_0$ and $\alpha_1$. (Compare with the worked example in Chapter 8.1.)

$\g_0=-\b$. After step 1, recompute $\g_1=A\x_1-\b$ and use $\d_1^\top A\d_1=10$.

$\alpha_0=-\g_0^\top\d_0/\d_0^\top A\d_0=3/2$, $\x_1=(\tfrac32,0)^\top$, $\g_1=(3,\tfrac32)^\top-(3,4)^\top=(0,-\tfrac52)^\top$ (and $\g_1^\top\d_0=0$). $\alpha_1=-\g_1^\top\d_1/10=-5/10=-\tfrac12$, so $\x_2=(1,1)^\top=\x^\star$. From $\x_0=\0$ the step lengths equal the expansion coefficients $a_j$: that is the $\sigma_k=\alpha_k$ step of the proof.

$A=\mathrm{diag}(1,2,4)$, $\b=(1,2,4)^\top$, $\x_0=\0$. Do exact line searches along $\e_1$ and then $\e_2$. What is $f(\x_2)-f^\star$?

For diagonal $A$ the axes are conjugate. By the Expanding Subspace Theorem, $\x_2$ minimizes $f$ over $\operatorname{span}\{\e_1,\e_2\}$, so its first two entries are already optimal and the third is untouched.

$\x^\star=(1,1,1)^\top$ and $\x_2=(1,1,0)^\top$. $f(\x_2)=\tfrac12(1+2)-(1+2)=-1.5$ and $f^\star=-\tfrac12\b^\top\x^\star=-3.5$, so the gap is $2$ (equivalently $\tfrac12\norm{\x_2-\x^\star}_A^2=\tfrac12\cdot4\cdot1=2$).

In the inductive step of Expanding Subspace (i), $\g_{k+1}^\top\d_i=\g_k^\top\d_i+\alpha_k\d_k^\top A\d_i$ for $i\lt k$. The induction hypothesis kills the first term. What kills the second?

Which property fails for the coordinate axes in practice problem 1?

Conjugacy: $\d_k^\top A\d_i=0$ for $i\ne k$. (The exact line search handles only the case $i=k$.)

  • Derive $\alpha_k$ by minimizing the parabola $\phi(\alpha)$, and remember $\alpha_k$ can be negative
  • Use $\g_{k+1}=\g_k+\alpha_kA\d_k$ (it avoids recomputing $A\x_{k+1}$ and drives the proof)
  • State Expanding Subspace precisely: $\x_k$ minimizes $f$ over $\x_0+\operatorname{span}\{\d_0,\dots,\d_{k-1}\}$
  • Check $\g_k^\top\d_i=0$ while computing by hand
  • Using inexact (Armijo/Wolfe) steps and expecting $n$-step termination: the proof needs $\g_{k+1}^\top\d_k=0$ exactly
  • Saying "$\x_k$ minimizes $f$ over $\mathcal B_k$" (forgetting the shift by $\x_0$)
  • Expecting finite termination in floating point for large $n$: round-off slowly destroys conjugacy
  1. Exact line search along $\d_k$ gives $\alpha_k=-\g_k^\top\d_k/\d_k^\top A\d_k$, $\g_{k+1}=\g_k+\alpha_kA\d_k$ and $\g_{k+1}^\top\d_k=0$.
  2. With $n$ conjugate directions, $\x_n=\x^\star$ from any start (Thm 5.1): in conjugate coordinates the problem is separable.
  3. Expanding Subspace (Thm 5.2): $\g_k\perp\d_0,\dots,\d_{k-1}$, so $\x_k$ minimizes $f$ (equivalently $\norm{\x-\x^\star}_A$) over $\x_0+\operatorname{span}\{\d_0,\dots,\d_{k-1}\}$.

After 3 steps of the conjugate direction method in $\R^{10}$, which must hold?

$\g_3=\0$
That happens at step $n=10$ (or earlier only by luck).
$\g_3^\top\d_0=\g_3^\top\d_1=\g_3^\top\d_2=0$
Part (i) with $k=3$.
$\g_3^\top\g_0=0$
Not for general conjugate directions. For CG it does hold (Chapter 8.3), but that needs more structure.

$\x_2$ from a conjugate direction run in $\R^3$ minimizes $f$ over…

$\operatorname{span}\{\d_0,\d_1\}$
The affine set must pass through the starting point.
all of $\R^3$
That's $\x_3$.
$\x_0+\operatorname{span}\{\d_0,\d_1\}$
The plane through $\x_0$ spanned by the first two directions.

Why does the order of the conjugate directions not matter for $\x_n=\x^\star$?

Each step adds exactly $\sigma_k\d_k$, the $k$-th coefficient of $\x^\star-\x_0$ in the conjugate basis, regardless of the other steps
The proof shows $\alpha_k=\sigma_k$; the sum of the same terms in any order is the same.
Because the directions are orthogonal
They needn't be orthogonal at all.
It does matter; you must start with the eigenvector of the largest eigenvalue
No ordering appears anywhere in the proof.

$f(\x)-f^\star$ equals…

$\tfrac12\norm{\x-\x^\star}^2$
Only when $A=I$. Which norm matches the curvature of $f$?
$\tfrac12\norm{\x-\x^\star}_A^2$
So minimizing $f$ over a set = finding the point closest to $\x^\star$ in the $A$-norm.
$\tfrac12\norm{\g(\x)}^2$
That's $\tfrac12\norm{\x-\x^\star}_{A^2}^2$, a different quantity.

The conjugate gradient method and why it is correct

CG builds conjugate directions on the fly: take the steepest descent direction and bend it just enough to be conjugate to the previous direction, $\d_k=-\g_k+\beta_k\d_{k-1}$. Conjugacy to all older directions then comes for free.

That makes CG cost about the same per step as gradient descent (one matrix–vector product, a few vectors of storage) while finishing in at most $n$ steps. It is the standard method for large sparse positive definite systems, and the CG Theorem's proof and a by-hand run are classic exam questions.

Steering a boat in a crosswind: you don't point straight at the target (the gradient); you keep a bit of your previous heading so you stop drifting back and forth.

Deriving the algorithm

Chapter 8.2 assumed someone handed us conjugate directions. CG chooses $\d_0=-\g_0$ (a steepest descent step) and then $$\d_k=-\g_k+\beta_k\d_{k-1},\qquad\beta_k\text{ chosen so that }\d_k^\top A\d_{k-1}=0.$$ Substituting: $0=-\g_k^\top A\d_{k-1}+\beta_k\,\d_{k-1}^\top A\d_{k-1}$, so $$\beta_k=\frac{\g_k^\top A\d_{k-1}}{\d_{k-1}^\top A\d_{k-1}}.$$ The step size is the exact line search from Chapter 8.2. The term $\beta_k\d_{k-1}$ is a memory of the previous heading, the very thing gradient descent lacks.

Given $\x_0$: $\g_0=A\x_0-\b$, $\d_0=-\g_0$. For $k=0,1,\dots$ while $\g_k\ne\0$: $$\alpha_k=-\frac{\g_k^\top\d_k}{\d_k^\top A\d_k},\quad \x_{k+1}=\x_k+\alpha_k\d_k,\quad \g_{k+1}=A\x_{k+1}-\b,$$ $$\beta_{k+1}=\frac{\g_{k+1}^\top A\d_k}{\d_k^\top A\d_k},\quad \d_{k+1}=-\g_{k+1}+\beta_{k+1}\d_k.$$

We only forced $\d_{k+1}$ to be conjugate to $\d_k$. Why would it be conjugate to $\d_{k-1},\d_{k-2},\dots$? That is the content of the CG Theorem, and its proof uses one more idea.

Krylov subspaces: what CG can "see"

After one step you've moved along $\g_0$, and the new gradient $\g_1=\g_0+\alpha_0A\d_0$ involves $A\g_0$. After two steps, the new gradient involves $A^2\g_0$. Each step applies $A$ once more.

$\mathcal K_k(\g_0)=\operatorname{span}\{\g_0,A\g_0,A^2\g_0,\dots,A^k\g_0\}$, the Krylov subspace of degree $k$ (at most $k+1$ dimensions). We write $\mathcal K_k$. Note $A\,\mathcal K_k\subseteq\mathcal K_{k+1}$. (This is [NW]'s $\mathcal K(\r_0;k)$, eq. (5.14), with $\r_0=\g_0$.)

Run CG from any $\x_0$ and suppose $\g_0,\dots,\g_k$ are all nonzero (the method has not stopped). Then

  1. $\operatorname{span}\{\g_0,\dots,\g_k\}=\mathcal K_k$;
  2. $\operatorname{span}\{\d_0,\dots,\d_k\}=\mathcal K_k$, and $\dim\mathcal K_k=k+1$;
  3. $\d_k^\top A\d_i=0$ for all $i\lt k$: all directions are mutually conjugate;
  4. $\g_k^\top\g_i=0$ for all $i\lt k$: all gradients are mutually orthogonal;
  5. $\g_k^\top\d_k=-\norm{\g_k}^2\lt0$: every $\d_k$ is a descent direction (and nonzero).

Hence CG is a conjugate direction method, and by Thm 5.1 it reaches $\x^\star$ in at most $n$ steps.

Read it before proving it. (c) is the surprise: one condition imposed, all conditions obtained. (d) gives a second view of termination: there's room for at most $n$ nonzero orthogonal vectors in $\R^n$, so some $\g_k$ with $k\le n$ must be $\0$. Compare gradient descent, which only guarantees $\g_{k+1}\perp\g_k$. (e) makes $\alpha_k\gt0$ and gives cheaper formulas.

Prove the CG Theorem by induction on $k$ (following [LD]). Write $\mathcal D_k=\operatorname{span}\{\d_0,\dots,\d_k\}$ and $\mathcal G_k=\operatorname{span}\{\g_0,\dots,\g_k\}$.

  1. Base $k=0$. $\mathcal G_0=\mathcal D_0=\operatorname{span}\{\g_0\}=\mathcal K_0$, one-dimensional since $\g_0\ne\0$. (c), (d) say nothing yet. (e): $\g_0^\top\d_0=-\norm{\g_0}^2$. Now assume (a)–(e) up to $k$ and $\g_{k+1}\ne\0$.

    Everything is checked directly from $\d_0=-\g_0$.

  2. Step 1: $\g_{k+1}\perp\mathcal D_k$. By (c), $\d_0,\dots,\d_k$ are nonzero and mutually conjugate, and CG used exact line searches along them. So $\x_0,\dots,\x_{k+1}$ are conjugate direction iterates, and Expanding Subspace (i) gives $\g_{k+1}^\top\d_i=0$ for all $i\le k$.

    This is where Chapter 8.2 plugs in. Everything else is bookkeeping about spans.

  3. Step 2: (a) at $k+1$. $\g_{k+1}=\g_k+\alpha_kA\d_k$ with $\g_k\in\mathcal K_k$ and $\d_k\in\mathcal K_k$, so $A\d_k\in\mathcal K_{k+1}$ and $\g_{k+1}\in\mathcal K_{k+1}$. Also $\g_{k+1}\notin\mathcal K_k$: otherwise $\g_{k+1}\in\mathcal D_k$ and Step 1 would make it orthogonal to itself, i.e. $\0$. So $\dim\mathcal G_{k+1}=k+2$, and $\mathcal G_{k+1}\subseteq\mathcal K_{k+1}$ with $\dim\mathcal K_{k+1}\le k+2$. Equal dimensions and inclusion give $\mathcal G_{k+1}=\mathcal K_{k+1}$.

    "Not in the old span" needs $\g_{k+1}\ne\0$; that's why the theorem assumes the method hasn't stopped.

  4. Step 3: (b) at $k+1$. $\d_{k+1}=-\g_{k+1}+\beta_{k+1}\d_k$, so $\mathcal D_{k+1}=\operatorname{span}(\mathcal D_k\cup\{\g_{k+1}\})=\operatorname{span}(\mathcal G_k\cup\{\g_{k+1}\})=\mathcal G_{k+1}=\mathcal K_{k+1}$.

    Adding $\g_{k+1}$ or adding $-\g_{k+1}+\beta\d_k$ to a span that already contains $\d_k$ gives the same span.

  5. Step 4: (e) at $k+1$. $\g_{k+1}^\top\d_{k+1}=-\norm{\g_{k+1}}^2+\beta_{k+1}\g_{k+1}^\top\d_k=-\norm{\g_{k+1}}^2$, since $\g_{k+1}^\top\d_k=0$ (Step 1). This is nonzero, so $\d_{k+1}\ne\0$.

    A nonzero direction is needed for it to count as a conjugate direction.

  6. Step 5 (the key step): (c) at $k+1$. For $i\le k$: $\d_{k+1}^\top A\d_i=-\g_{k+1}^\top A\d_i+\beta_{k+1}\d_k^\top A\d_i$. If $i=k$, this is 0 by the choice of $\beta_{k+1}$. If $i\lt k$: the second term is 0 by (c) at stage $k$; for the first, $\d_i\in\mathcal K_i$, so $A\d_i\in\mathcal K_{i+1}\subseteq\mathcal K_k=\mathcal D_k$, and $\g_{k+1}\perp\mathcal D_k$ by Step 1.

    Conjugacy to old directions is free because $A\d_i$ stays inside the Krylov space that the new gradient is already orthogonal to.

  7. Step 6: (d) at $k+1$. For $i\le k$, $\g_i\in\mathcal G_k=\mathcal D_k$, and $\g_{k+1}\perp\mathcal D_k$ (Step 1), so $\g_{k+1}^\top\g_i=0$. Induction complete. Since at most $n$ nonzero conjugate directions exist, some $\g_k$ with $k\le n$ is $\0$, i.e. $\x_k=\x^\star$. $\blacksquare$

    Skeleton to memorize: (1) Expanding Subspace gives $\g_{k+1}\perp\mathcal D_k$; (2) $\g_{k+1}=\g_k+\alpha A\d_k$ grows the Krylov space by one; (3) $A\mathcal K_i\subseteq\mathcal K_{i+1}$ gives conjugacy to old directions.

While $\g_k\ne\0$: $\qquad\alpha_k=\dfrac{\g_k^\top\g_k}{\d_k^\top A\d_k}\gt0,\qquad\beta_{k+1}=\dfrac{\g_{k+1}^\top\g_{k+1}}{\g_k^\top\g_k}.$

Proof (examinable). For $\alpha_k$: $-\g_k^\top\d_k=\norm{\g_k}^2$ by (e). For $\beta_{k+1}$: from $\g_{k+1}=\g_k+\alpha_kA\d_k$, $A\d_k=(\g_{k+1}-\g_k)/\alpha_k$. The numerator of the preliminary $\beta_{k+1}$ is $\g_{k+1}^\top A\d_k=(\norm{\g_{k+1}}^2-\g_{k+1}^\top\g_k)/\alpha_k=\norm{\g_{k+1}}^2/\alpha_k$ by (d). The denominator is $\d_k^\top A\d_k=\norm{\g_k}^2/\alpha_k$ by the new $\alpha_k$ formula. Divide; $\alpha_k$ cancels. $\blacksquare$

If an exam asks you to derive CG, start from the preliminary $\beta$ and the line-search $\alpha$, then simplify using (d) and (e). The cheap formulas are consequences of the theorem, not definitions.

$\g_0\gets A\x_0-\b$; $\d_0\gets-\g_0$; $k\gets0$. While $\norm{\g_k}\gt\text{tol}$:

  1. $\mathbf w\gets A\d_k$ (the only matrix–vector product)
  2. $\alpha_k\gets\norm{\g_k}^2/(\d_k^\top\mathbf w)$
  3. $\x_{k+1}\gets\x_k+\alpha_k\d_k$; $\ \g_{k+1}\gets\g_k+\alpha_k\mathbf w$
  4. $\beta_{k+1}\gets\norm{\g_{k+1}}^2/\norm{\g_k}^2$; $\ \d_{k+1}\gets-\g_{k+1}+\beta_{k+1}\d_k$; $\ k\gets k+1$
MethodWork per iterationStorage
CG1 product $A\d$, 2 dot products, 3 vector updates4 vectors ($\x,\g,\d,A\d$)
Gradient descent, exact step1 product $A\g$, 2 dot products, 2 vector updates3 vectors
Cholesky (direct solve)about $n^3/3$ operations, once$n^2$ numbers for dense $A$

CG needs $A$ only as a function $\v\mapsto A\v$. If $A$ is sparse with a handful of nonzeros per row, each iteration costs $O(n)$, which is why CG can handle systems with millions of unknowns that a direct method couldn't even store.

Try it

Gradient descent (orange, exact steps) and CG (blue) start from the same dark point. Drag it anywhere: CG lands on the star in exactly 2 steps every time, while gradient descent zig-zags. Push $\kappa$ to 60: the zig-zag gets worse, CG doesn't care. Switch to "stretched view": the contours become circles, the two CG steps become perpendicular (conjugate = orthogonal after stretching), and the second step points straight at the centre. Then choose "fit the line": from $(0,0)$ gradient descent is lucky (about 5 steps to $10^{-6}$), but drag the start to $(1,2)$ and it needs about 50; CG needs 2 from anywhere.

Run CG by hand on $A=\begin{pmatrix}4&1\\1&3\end{pmatrix}$, $\b=(1,2)^\top$, from $\x_0=\0$, checking the theorem as you go.

  1. $\g_0=-\b=(-1,-2)^\top$, $\d_0=(1,2)^\top$, $A\d_0=(6,7)^\top$, $\d_0^\top A\d_0=20$, $\norm{\g_0}^2=5$. So $\alpha_0=\tfrac5{20}=\tfrac14$ and $\x_1=(\tfrac14,\tfrac12)^\top$.

    The first CG step is exactly a steepest descent step with exact line search.

  2. $\g_1=\g_0+\tfrac14A\d_0=(\tfrac12,-\tfrac14)^\top$. Check (d): $\g_1^\top\g_0=-\tfrac12+\tfrac12=0$. ✓

    Use the update $\g_{k+1}=\g_k+\alpha_kA\d_k$: it reuses $A\d_0$.

  3. $\beta_1=\norm{\g_1}^2/\norm{\g_0}^2=\tfrac{5/16}{5}=\tfrac1{16}$, and $\d_1=-\g_1+\tfrac1{16}\d_0=(-\tfrac7{16},\tfrac38)^\top$. Check (c): $\d_1^\top A\d_0=-\tfrac{42}{16}+\tfrac{21}8=0$. ✓

    A small $\beta$ means only a slight bend away from $-\g_1$.

  4. $A\d_1=(-\tfrac{11}8,\tfrac{11}{16})^\top$, $\d_1^\top A\d_1=\tfrac{77}{128}+\tfrac{33}{128}=\tfrac{55}{64}$, so $\alpha_1=\tfrac{5/16}{55/64}=\tfrac4{11}$.

    Practical $\alpha$: numerator $\norm{\g_1}^2$.

  5. $\x_2=(\tfrac14,\tfrac12)^\top+\tfrac4{11}(-\tfrac7{16},\tfrac38)^\top=(\tfrac1{11},\tfrac7{11})^\top=\x^\star$, and $\g_2=\0$. Stop after $n=2$ steps.

    Gradient descent from the same start is still about $2.6\times10^{-6}$ away after 10 steps.

A $3\times3$ run: $A=\begin{pmatrix}2&-1&0\\-1&2&-1\\0&-1&2\end{pmatrix}$, $\b=(1,0,0)^\top$, $\x_0=\0$. Watch the Krylov subspaces grow.

  1. $\g_0=(-1,0,0)^\top$, $\d_0=(1,0,0)^\top$, $A\d_0=(2,-1,0)^\top$, $\d_0^\top A\d_0=2$, $\norm{\g_0}^2=1$: $\alpha_0=\tfrac12$, $\x_1=(\tfrac12,0,0)^\top$, $\g_1=\g_0+\tfrac12A\d_0=(0,-\tfrac12,0)^\top$.

    $\mathcal K_0=\operatorname{span}\{\e_1\}$.

  2. $\beta_1=\tfrac{1/4}{1}=\tfrac14$, $\d_1=-\g_1+\tfrac14\d_0=(\tfrac14,\tfrac12,0)^\top$, $A\d_1=(0,\tfrac34,-\tfrac12)^\top$, $\d_1^\top A\d_1=\tfrac38$: $\alpha_1=\tfrac{1/4}{3/8}=\tfrac23$, $\x_2=(\tfrac23,\tfrac13,0)^\top$, $\g_2=\g_1+\tfrac23A\d_1=(0,0,-\tfrac13)^\top$.

    Because $A$ is tridiagonal, each power of $A$ reaches one more coordinate: $\mathcal K_1=\operatorname{span}\{\e_1,\e_2\}$, and indeed $\x_2,\d_1$ have zero third entry.

  3. $\beta_2=\tfrac{1/9}{1/4}=\tfrac49$, $\d_2=-\g_2+\tfrac49\d_1=(\tfrac19,\tfrac29,\tfrac13)^\top$, $A\d_2=(0,0,\tfrac49)^\top$, $\d_2^\top A\d_2=\tfrac4{27}$: $\alpha_2=\tfrac{1/9}{4/27}=\tfrac34$.

    The gradients $\g_0,\g_1,\g_2$ point along $\e_1,\e_2,\e_3$: mutually orthogonal, as (d) promises.

  4. $\x_3=(\tfrac23,\tfrac13,0)^\top+\tfrac34(\tfrac19,\tfrac29,\tfrac13)^\top=(\tfrac34,\tfrac12,\tfrac14)^\top$. Check: $A\x_3=(\tfrac32-\tfrac12,\,-\tfrac34+1-\tfrac14,\,-\tfrac12+\tfrac12)^\top=(1,0,0)^\top=\b$. ✓

    Exactly $n=3$ steps. The CG directions here are parallel to the $A$-Gram–Schmidt vectors of $\e_1,\e_2,\e_3$ from Chapter 8.1, because the gradients are multiples of $\e_1,\e_2,\e_3$: CG is $A$-Gram–Schmidt applied to the gradients, done with a two-term recurrence.

Try it

Step through CG on the $3\times3$ system and compare with the worked example (fractions are shown exactly). The grids check the theorem live: $\d_i^\top A\d_j$ and $\g_i^\top\g_j$ off the diagonal are green zeros, and the Krylov line confirms (b). Then pick $\b=(1,0,1)$: CG finishes after only 2 steps. (That vector involves only two of $A$'s three eigenvectors; Part 9 explains why that matters.)

Why only one $\beta$? Symmetry

Gram–Schmidt must correct a new vector against every old direction. In CG all corrections except the last vanish (Step 5), because $\g_{k+1}^\top A\d_i=(A\g_{k+1})^\top\d_i$ needs $A=A^\top$ to move $A$ across, and because the Krylov structure keeps $A\d_i$ inside the explored space. For non-symmetric $A$ there is no such short recurrence, and Krylov methods for those systems (such as GMRES) must store all old directions.

A teaser for Part 9. Finite termination is an exact-arithmetic statement, and for $n$ in the millions you never run $n$ steps anyway. CG is used as an iterative method that is often accurate after far fewer than $n$ steps. How few depends on the eigenvalues of $A$, and Part 9 makes this precise.

Go deeper: CG in floating point, and nonlinear CG

Round-off. The proof enforces conjugacy only with the previous direction; global conjugacy and orthogonality are inherited through the induction, and in floating point those inherited properties decay first. In the book's lab ($n=40$), with $\kappa=10^2$ CG needs 54 iterations (not 40) to reach a relative $A$-norm error of $10^{-10}$, with $\kappa=10^4$ it needs 105, and with $\kappa=10^6$ it hasn't got there after 160. The error still decreases monotonically, because each step is still an exact minimization along $\d_k$.

Nonlinear CG (not a lecture topic). The practical formulas use only gradients, so they make sense for any smooth $f$: use a line search for $\alpha_k$ and $\beta^{\mathrm{FR}}_{k+1}=\norm{\g_{k+1}}^2/\norm{\g_k}^2$ (Fletcher–Reeves) or $\beta^{\mathrm{PR}}_{k+1}=\g_{k+1}^\top(\g_{k+1}-\g_k)/\norm{\g_k}^2$ (Polak–Ribière). On a quadratic with exact line searches they coincide, by part (d). See [NW] §5.2 and [FR] §4.1.

A harder exercise. Show that $f(\x_{k+1})\le f(\x_k-t\g_k)$ for every $t$: CG is never worse than one steepest descent step from the same point. (Hint: $\x_{k+1}$ minimizes $f$ over $\x_0+\mathcal K_k$, and both $\x_k$ and $\g_k$ lie in the right sets by (a)–(b).)

Run CG on $A=\begin{pmatrix}2&1\\1&2\end{pmatrix}$, $\b=(1,0)^\top$ from $\x_0=\0$. Give $\alpha_0$, $\beta_1$ and $\x_2$.

$\g_0=(-1,0)^\top$, $\d_0=(1,0)^\top$, $A\d_0=(2,1)^\top$. Then $\g_1=\g_0+\alpha_0A\d_0$.

$\alpha_0=\tfrac12$, $\x_1=(\tfrac12,0)^\top$, $\g_1=(0,\tfrac12)^\top$ ($\g_1^\top\g_0=0$ ✓). $\beta_1=\tfrac{1/4}{1}=\tfrac14$, $\d_1=(\tfrac14,-\tfrac12)^\top$ ($\d_1^\top A\d_0=\tfrac12-\tfrac12=0$ ✓). $A\d_1=(0,-\tfrac34)^\top$, $\d_1^\top A\d_1=\tfrac38$, $\alpha_1=\tfrac{1/4}{3/8}=\tfrac23$. $\x_2=(\tfrac12+\tfrac16,\,-\tfrac13)^\top=(\tfrac23,-\tfrac13)^\top$; check $A\x_2=(1,0)^\top=\b$.

Fit the line: $A=\begin{pmatrix}28&12\\12&6\end{pmatrix}$, $\b=(46,20)^\top$, start $\z_0=(0,0)^\top$. Compute the first CG step: $\alpha_0$ and $\z_1=(w_1,c_1)$. (Give 4 significant digits.)

$\g_0=-\b$, $\d_0=\b$, $\norm{\g_0}^2=46^2+20^2$. Compute $A\d_0$ and $\d_0^\top A\d_0$.

$\norm{\g_0}^2=2116+400=2516$. $A\d_0=(28\cdot46+12\cdot20,\ 12\cdot46+6\cdot20)^\top=(1528,672)^\top$, $\d_0^\top A\d_0=46\cdot1528+20\cdot672=83728$. So $\alpha_0=2516/83728\approx0.03005$ and $\z_1=\alpha_0(46,20)\approx(1.3823,\,0.6010)$. One more CG step lands on $(1.5,\tfrac13)$ exactly.

Same tridiagonal $A$ as the $3\times3$ worked example, but $\b=(1,0,1)^\top$ and $\x_0=\0$. Compute $\alpha_0$, $\beta_1$, $\alpha_1$. (Then check that $\x_2$ already solves the system.)

$\g_0=(-1,0,-1)^\top$, $\d_0=(1,0,1)^\top$, $A\d_0=(2,-2,2)^\top$.

$\norm{\g_0}^2=2$, $\d_0^\top A\d_0=4$, $\alpha_0=\tfrac12$, $\x_1=(\tfrac12,0,\tfrac12)^\top$, $\g_1=\g_0+\tfrac12A\d_0=(0,-1,0)^\top$. $\beta_1=\tfrac12$, $\d_1=(0,1,0)^\top+\tfrac12(1,0,1)^\top=(\tfrac12,1,\tfrac12)^\top$, $A\d_1=(0,1,0)^\top$, $\d_1^\top A\d_1=1$, $\alpha_1=\tfrac11=1$. $\x_2=(1,1,1)^\top$ and $A\x_2=(1,0,1)^\top=\b$: done in 2 steps, fewer than $n=3$.

In deriving $\beta_{k+1}=\norm{\g_{k+1}}^2/\norm{\g_k}^2$ from $\beta_{k+1}=\g_{k+1}^\top A\d_k/\d_k^\top A\d_k$, the numerator becomes $(\norm{\g_{k+1}}^2-\g_{k+1}^\top\g_k)/\alpha_k$. Which fact removes the second term?

Which part of the CG Theorem is about gradients being perpendicular to each other?

Part (d), mutual orthogonality of the gradients, gives $\g_{k+1}^\top\g_k=0$. (Part (e) is then used for the denominator.)

  • Use the practical formulas and the update $\g_{k+1}=\g_k+\alpha_kA\d_k$: one product $A\d_k$ per step
  • Check $\g_1^\top\g_0=0$ and $\d_1^\top A\d_0=0$ during hand computations
  • Reproduce the proof via its skeleton: Expanding Subspace → Krylov growth → $A\mathcal K_i\subseteq\mathcal K_{i+1}$
  • Remember the first CG step is a steepest descent step
  • Writing $\beta=\norm{\g_{k+1}}^2/\norm{\g_k}^2$ as a definition when asked to derive CG
  • Mixing sign conventions: with $\r=\b-A\x$ the direction update reads $\d_{k+1}=\r_{k+1}+\beta\d_k$
  • Running CG on an indefinite $A$ ($\d^\top A\d$ can be $\le0$ and the method breaks)
  • Forgetting the hypothesis "$\g_{k+1}\ne\0$" in Step 2 of the proof
  1. CG: $\d_0=-\g_0$, $\alpha_k=\norm{\g_k}^2/\d_k^\top A\d_k$, $\g_{k+1}=\g_k+\alpha_kA\d_k$, $\beta_{k+1}=\norm{\g_{k+1}}^2/\norm{\g_k}^2$, $\d_{k+1}=-\g_{k+1}+\beta_{k+1}\d_k$.
  2. CG Theorem: gradients and directions span the Krylov spaces $\mathcal K_k(\g_0)$, directions are mutually conjugate, gradients mutually orthogonal, so CG terminates in at most $n$ steps.
  3. Per step: one matrix–vector product and $O(n)$ extra work and storage, about the cost of gradient descent.

CG enforces $\d_{k+1}^\top A\d_k=0$ only. Why is $\d_{k+1}^\top A\d_i=0$ for $i\lt k$ too?

Because $\beta_{k+1}$ is chosen to cancel all old components
One scalar can only enforce one condition.
Because $A\d_i\in\mathcal K_{i+1}\subseteq\mathcal D_k$ and $\g_{k+1}\perp\mathcal D_k$, while $\d_k^\top A\d_i=0$ already
Step 5 of the proof.
Because the gradients are orthogonal
That's another conclusion of the theorem, not the mechanism for conjugacy.

CG on $A\succ0$ with $n=1000$, in exact arithmetic, starting anywhere. Which is guaranteed?

$\x_2=\x^\star$
Only for $n=2$ (or special data).
$\g_k=\0$ for some $k\le1000$
At most $n$ nonzero mutually orthogonal gradients fit in $\R^{1000}$.
$f$ decreases by the same amount every step
Progress can be very uneven; Part 9 studies how fast it goes.

[NW] writes $\r_k=A\x_k-\b$ and Alg 5.2 has $\p_{k+1}=-\r_{k+1}+\beta_{k+1}\p_k$. A text that defines $\r_k=\b-A\x_k$ would write…

$\p_{k+1}=-\r_{k+1}+\beta_{k+1}\p_k$ with the same $\beta$
The sign of $\r$ flipped, so something in the update must flip too.
$\p_{k+1}=\r_{k+1}+\beta_{k+1}\p_k$, with $\beta_{k+1}=\norm{\r_{k+1}}^2/\norm{\r_k}^2$ unchanged
$\b-A\x=-\g$; norms don't see the sign.
$\p_{k+1}=\r_{k+1}-\beta_{k+1}\p_k$
$\beta$ is a ratio of squared norms, so its sign doesn't change with the convention.

Which property of $A$ makes the one-term recurrence possible (beyond $\d^\top A\d\gt0$)?

$A$ is diagonal
CG works for any symmetric positive definite $A$.
$A$ is sparse
Sparsity makes each step cheap but isn't needed for correctness.
$A$ is symmetric
It lets $A$ move across inner products; without it methods like GMRES store all directions.

Gradient descent with exact line search on a quadratic, and CG, both satisfy $\g_{k+1}^\top\g_k=0$. What does CG have that gradient descent lacks?

Exact line searches
Both use them.
$\g_{k+1}\perp\g_i$ for all $i\le k$, so it can never revisit an old direction
In 2-D, gradient descent's $\g_{k+2}$ is parallel to $\g_k$: the zig-zag.
A larger step size
Step size isn't the difference; the directions are.

For fit the line ($\kappa\approx46.1$), the classical bound for exact-step gradient descent is $f(\x_{k+1})-f^\star\le\left(\frac{\kappa-1}{\kappa+1}\right)^2(f(\x_k)-f^\star)$. Roughly what factor is that per step?

0.04
That would be fast. Compute $(45.1/47.1)^2$.
0.92
$(45.14/47.14)^2\approx0.917$: only about 8% guaranteed progress per step, whereas CG finishes in 2.
0.50
Compute $(\kappa-1)/(\kappa+1)$ first.

Proving Expanding Subspace (ii) from (i) uses which fact from Part 4?

A stationary point of a convex function is a global minimizer
$h(\y)=f(\x_0+D_k\y)$ is convex with $\grad h=D_k^\top\g_k=\0$ at $\x_k$.
Every local minimizer of any function is global
False in general; convexity is what makes it true.
A positive definite Hessian implies the gradient is zero
The Hessian says nothing about where the gradient vanishes.

For $A=V\Lambda V^\top\succ0$, which matrix turns $A$-conjugacy into ordinary orthogonality?

$V^\top$ (rotate to the eigenbasis)
A rotation preserves ordinary angles, so it can't create perpendicularity.
$A^{1/2}=V\Lambda^{1/2}V^\top$ (or $\Lambda^{1/2}V^\top$)
$\u^\top A\v=(A^{1/2}\u)^\top(A^{1/2}\v)$.
$A^{-1}$
Then $(A^{-1}\u)^\top(A^{-1}\v)=\u^\top A^{-2}\v$, the wrong matrix.

In CG, $\x_k-\x_0$ always lies in…

$\operatorname{span}\{\g_k\}$
The iterate collects all previous steps, not just one.
$\mathcal K_{k-1}(\g_0)=\operatorname{span}\{\g_0,\dots,A^{k-1}\g_0\}$
$\x_k-\x_0=\sum_{i\lt k}\alpha_i\d_i$ and $\operatorname{span}\{\d_0,\dots,\d_{k-1}\}=\mathcal K_{k-1}$. This is Part 9's starting point.
the eigenvector of the smallest eigenvalue
Nothing forces the iterates onto an eigenvector.

Lecture 13 (24 Sep): the polynomial interpretation of the conjugate gradient method. After $k$ steps, CG's error is a polynomial in $A$ applied to the starting error, and CG picks the best such polynomial. From that one idea come all of CG's convergence results: few distinct eigenvalues means few steps, clusters are almost as good, CG beats every gradient method, and the famous $\sqrt\kappa$ rate.

You need: Part 8 (CG algorithm, Krylov subspaces, Expanding Subspace Theorem), Part 5 (the steepest descent rate $\frac{\kappa-1}{\kappa+1}$ on quadratics), Part 0b (eigenvalues and the spectral theorem).

From Krylov subspaces to polynomials: CG picks the best one

After $k$ steps, CG's error equals a polynomial in the matrix $A$ applied to the starting error, and CG automatically chooses the best such polynomial.

"CG finishes in at most $n$ steps" is useless when $n=10^6$. This viewpoint tells you how good $\x_k$ is for small $k$, and it is the basis of several short exam proofs.

A sound engineer with a $k$-band equaliser: the error is a mix of "frequencies" (eigenvalues), and CG sets the $k$ knobs to silence that mix as well as possible. With only a few distinct frequencies, it can silence them completely.

Recall the setting from Part 8. We minimize $f(\x)=\tfrac12\x^\top A\x-\b^\top\x$ with $A$ symmetric positive definite ($A\succ0$), so the minimizer is $\x^\star=A^{-1}\b$ and the gradient is $\g_k=\grad f(\x_k)=A\x_k-\b$. Part 8 gave us three facts about CG:

  • CG Theorem: the directions span a Krylov subspace, $\operatorname{span}\{\d_0,\dots,\d_k\}=\mathcal K_k(\g_0):=\operatorname{span}\{\g_0,A\g_0,\dots,A^k\g_0\}$.
  • Expanding Subspace Theorem: $\x_{k+1}$ minimizes $f$ over the whole set $\x_0+\operatorname{span}\{\d_0,\dots,\d_k\}$, not just along the last direction.
  • Error identity: $f(\x)-f(\x^\star)=\tfrac12(\x-\x^\star)^\top A(\x-\x^\star)$. So minimizing $f$ is the same as minimizing the distance to $\x^\star$ measured in the $A$-norm below.

Finite termination in $n$ steps follows. The question this part answers is the practical one: how small is the error after $k\ll n$ steps?

Notation for this part
$\e_k=\x_k-\x^\star$
the error after $k$ steps
$\norm{\z}_A=\sqrt{\z^\top A\z}$
the $A$-norm ("energy norm"); $f(\x)-f^\star=\tfrac12\norm{\x-\x^\star}_A^2$
$0<\lambda_1\le\dots\le\lambda_n$
eigenvalues of $A$, with orthonormal eigenvectors $\v_1,\dots,\v_n$; $m=\lambda_1$, $L=\lambda_n$, $\kappa=L/m$
$\xi_i$
coordinates of the starting error in the eigenbasis: $\e_0=\sum_i\xi_i\v_i$
$\mathcal P_k$
real polynomials of degree at most $k$
$\mathcal Q_k$
residual polynomials: those $Q\in\mathcal P_k$ with $Q(0)=1$

Plugging a matrix into a polynomial

For a polynomial $P(t)=\gamma_0+\gamma_1t+\dots+\gamma_kt^k$, define $P(A)=\gamma_0I+\gamma_1A+\dots+\gamma_kA^k$. If $A\v=\lambda\v$, then $A^2\v=A(\lambda\v)=\lambda^2\v$, and in general $A^j\v=\lambda^j\v$. Adding up:

$P(A)\v_i=P(\lambda_i)\,\v_i$ for every eigenpair $(\lambda_i,\v_i)$. So in the eigenbasis, $P(A)$ simply multiplies the $i$-th coordinate by the number $P(\lambda_i)$.

Example: $A=\begin{pmatrix}1&0\\0&4\end{pmatrix}$ and $P(t)=1-\tfrac t2$ give $P(A)=I-\tfrac12A=\begin{pmatrix}1/2&0\\0&-1\end{pmatrix}$: the diagonal entries are $P(1)$ and $P(4)$.

Step 1: CG's iterates are "$\x_0$ + a polynomial in $A$ times $\g_0$"

Since $\x_{k+1}=\x_0+\alpha_0\d_0+\dots+\alpha_k\d_k$ and the $\d_i$ span $\mathcal K_k(\g_0)$, there are numbers $\gamma_0,\dots,\gamma_k$ with $$\x_{k+1}=\x_0+\gamma_0\g_0+\gamma_1A\g_0+\dots+\gamma_kA^k\g_0=\x_0+P_k(A)\,\g_0,\qquad P_k\in\mathcal P_k .$$ This is [NW] (5.25) and [LD] §9.4 (24).

Step 2: the error is $Q(A)\e_0$ with $Q(0)=1$

Because $\b=A\x^\star$, the first gradient is $\g_0=A\x_0-A\x^\star=A\e_0$. Subtract $\x^\star$ from Step 1: $$\e_{k+1}=\e_0+P_k(A)A\e_0=\big[I+AP_k(A)\big]\e_0=Q_{k+1}(A)\,\e_0,\qquad Q_{k+1}(t):=1+t\,P_k(t).$$ $Q_{k+1}$ has degree at most $k+1$ and $Q_{k+1}(0)=1$. Conversely, any $Q$ of degree $\le k+1$ with $Q(0)=1$ arises this way: $Q(t)-1$ vanishes at $t=0$, so $P(t)=(Q(t)-1)/t$ is a polynomial of degree $\le k$.

$\mathcal Q_k=\{Q\in\mathcal P_k:\ Q(0)=1\}$. The map $P\mapsto Q(t)=1+tP(t)$ is a one-to-one correspondence between $\mathcal P_k$ and $\mathcal Q_{k+1}$. After $k$ steps of CG, $\e_k=Q_k(A)\e_0$ for some $Q_k\in\mathcal Q_k$.

Careful with indices: the point $\x_{k+1}$ uses a "position" polynomial $P_k$ of degree $k$ but an error polynomial $Q_{k+1}$ of degree $k+1$. Off-by-one slips here are the most common mistake in this topic. Also, the constraint is $Q(0)=1$, not "$Q$ is monic" (leading coefficient 1).
Gradient descent is a polynomial too

Take gradient descent with any step sizes: $\y_0=\x_0$, $\y_{j+1}=\y_j-h_j\grad f(\y_j)$. Since $\grad f(\y_j)=A(\y_j-\x^\star)$, $$\y_{j+1}-\x^\star=(I-h_jA)(\y_j-\x^\star)\quad\Longrightarrow\quad \y_k-\x^\star=Q(A)\e_0,\qquad Q(t)=\prod_{j=0}^{k-1}(1-h_jt).$$ This $Q$ has degree $k$ and $Q(0)=1$. Each step size $h_j$ places a root of $Q$ at $t=1/h_j$. A constant step $h$ gives $Q(t)=(1-ht)^k$: all $k$ roots stacked at the same place. So gradient descent, CG and many other methods all produce errors of the form $Q(A)\e_0$. What differs is which $Q$.

The $(k+1)$-st CG iterate satisfies $$\norm{\e_{k+1}}_A=\min_{P\in\mathcal P_k}\norm{\x_0+P(A)\g_0-\x^\star}_A=\min_{Q\in\mathcal Q_{k+1}}\norm{Q(A)\e_0}_A .$$ Among all points of $\x_0+\mathcal K_k(\g_0)$, the CG iterate is the closest to $\x^\star$ in the $A$-norm, and therefore has the smallest value of $f$.

Proof. The set $\{\x_0+P(A)\g_0:P\in\mathcal P_k\}$ is exactly $\x_0+\mathcal K_k(\g_0)=\x_0+\operatorname{span}\{\d_0,\dots,\d_k\}$. By the Expanding Subspace Theorem, $\x_{k+1}$ minimizes $f$ over this set, hence minimizes $f-f^\star=\tfrac12\norm{\cdot-\x^\star}_A^2$. That is the first equality. The second follows from $\x_0+P(A)\g_0-\x^\star=[I+AP(A)]\e_0=Q(A)\e_0$ and the correspondence $P\leftrightarrow Q$. $\blacksquare$

Step 3: eigenvalues turn the matrix problem into a scalar one

Expand the starting error in the eigenbasis, $\e_0=\sum_i\xi_i\v_i$. By the lemma, $Q(A)\e_0=\sum_iQ(\lambda_i)\xi_i\v_i$. For any coefficients $c_i$, $$\Big\lVert\sum_ic_i\v_i\Big\rVert_A^2=\sum_{i,j}c_ic_j\,\v_i^\top A\v_j=\sum_i\lambda_ic_i^2,$$ because $\v_i^\top A\v_j=\lambda_j\v_i^\top\v_j$ is $0$ for $i\ne j$ and $\lambda_i$ for $i=j$. Therefore $$\norm{Q(A)\e_0}_A^2=\sum_{i=1}^n\lambda_i\,Q(\lambda_i)^2\,\xi_i^2,\qquad \norm{\e_0}_A^2=\sum_{i=1}^n\lambda_i\,\xi_i^2 .$$ Read it like this: the error component along $\v_i$ gets multiplied by the number $Q(\lambda_i)$.

For every $k\ge0$ and every $Q\in\mathcal Q_{k+1}$, $$\norm{\e_{k+1}}_A^2\le\max_{1\le i\le n}Q(\lambda_i)^2\;\norm{\e_0}_A^2 .$$ Hence, after $k$ steps, $\displaystyle\norm{\e_k}_A\le\Big(\min_{Q\in\mathcal Q_k}\ \max_{i}\,|Q(\lambda_i)|\Big)\norm{\e_0}_A$.

Prove the min–max bound from the eigen-expansion formula. (A classic short exam proof.)

  1. Fix any $Q\in\mathcal Q_{k+1}$. By the optimality theorem, $\norm{\e_{k+1}}_A^2\le\norm{Q(A)\e_0}_A^2$.

    CG achieves the minimum over all of $\mathcal Q_{k+1}$, so any particular $Q$ can only do worse or equal.

  2. Expand: $\norm{Q(A)\e_0}_A^2=\sum_i\lambda_iQ(\lambda_i)^2\xi_i^2$.

    This is the eigen-expansion formula of Step 3.

  3. $\sum_i\lambda_iQ(\lambda_i)^2\xi_i^2\le\big(\max_jQ(\lambda_j)^2\big)\sum_i\lambda_i\xi_i^2$.

    Every weight $\lambda_i\xi_i^2$ is $\ge0$ (here $A\succ0$ is used), so replacing each $Q(\lambda_i)^2$ by the largest one can only increase the sum.

  4. $\sum_i\lambda_i\xi_i^2=\norm{\e_0}_A^2$, which gives the bound. The left side does not depend on $Q$, so we may take the minimum over $Q$ on the right; square roots give the second form.

    The bound holds for every $Q$ simultaneously, so it holds for the best one.

To bound CG after $k$ steps, design a polynomial of degree $\le k$ with $Q(0)=1$ that is small at the eigenvalues of $A$. It may do anything between the eigenvalues. A degree-$k$ polynomial has $k$ roots to spend: put them where the eigenvalues are. The bound only sees where the eigenvalues sit, not how many times each is repeated and not $\e_0$. CG plays this game automatically, and plays it at least as well as you, because it minimizes the actual weighted sum rather than the worst case.

Try it

The orange dots are six eigenvalues; $\e_0$ has equal weight on each. Your polynomial is $Q(t)=\prod_j(1-t/r_j)$ with roots $r_j$ at the purple diamonds (drag them). That is exactly gradient descent with step sizes $h_j=1/r_j$. Try to beat CG's error ratio at degree 2 or 3: you can't. Then press "Use CG's roots" to see the polynomial CG chose, and set the degree to 6 with "Roots on eigenvalues".

Let $A=\begin{pmatrix}1&0\\0&4\end{pmatrix}$, $\b=\0$ (so $\x^\star=\0$) and $\x_0=(1,1)^\top$. Run one CG step, find $Q_1$, and compare the actual error with the min–max bound.

  1. $\g_0=A\x_0=(1,4)^\top$, $\d_0=-\g_0$, $\alpha_0=\dfrac{\norm{\g_0}^2}{\d_0^\top A\d_0}=\dfrac{1+16}{1+64}=\dfrac{17}{65}$.

    CG's first step is steepest descent with exact line search.

  2. $\x_1=\x_0-\tfrac{17}{65}\g_0=\big(\tfrac{48}{65},-\tfrac{3}{65}\big)^\top=\e_1$. And $Q_1(t)=1-\tfrac{17}{65}t$ gives $Q_1(1)=\tfrac{48}{65}$, $Q_1(4)=-\tfrac{3}{65}$: exactly the two components of $\e_1$.

    $\e_0=(1,1)$ means $\xi_1=\xi_2=1$ (the eigenvectors are the coordinate axes), so component $i$ of $\e_1$ is $Q_1(\lambda_i)\cdot1$.

  3. $\norm{\e_1}_A^2=1\cdot\big(\tfrac{48}{65}\big)^2+4\cdot\big(\tfrac{3}{65}\big)^2=\tfrac{2340}{4225}\approx0.5538$ and $\norm{\e_0}_A^2=1+4=5$, so $\norm{\e_1}_A/\norm{\e_0}_A\approx0.333$.

    The weighted-sum formula with weights $\lambda_i\xi_i^2$.

  4. The min–max bound with $Q_1$ gives $\max(\tfrac{48}{65},\tfrac{3}{65})\approx0.738$. The best degree-1 worst case on $\{1,4\}$ is $Q(t)=1-\tfrac25t$, with $|Q(1)|=|Q(4)|=0.6=\frac{\kappa-1}{\kappa+1}$. CG's actual $0.333$ beats both.

    The bound must hold for every $\e_0$; CG adapts to this particular $\e_0$, whose larger weight sits on $\lambda=4$, so it pushes $Q$ close to $0$ there.

  5. At step 2, $Q_2(t)=(1-t)(1-\tfrac t4)$ vanishes at both eigenvalues, so $\e_2=\0$.

    A degree-2 polynomial with $Q(0)=1$ can have roots at both eigenvalues; CG's optimal choice must achieve error $0$, matching finite termination for $n=2$.

$A=\operatorname{diag}(1,3)$, $\b=\0$, $\x_0=(1,1)^\top$. After one CG step, $\e_1=Q_1(A)\e_0$. Find $Q_1(1)$ and $Q_1(3)$.

$Q_1(t)=1-\alpha_0t$ with $\alpha_0=\norm{\g_0}^2/\g_0^\top A\g_0$ and $\g_0=(1,3)$.

$\alpha_0=\frac{1+9}{1+27}=\frac{5}{14}$, so $Q_1(t)=1-\frac{5}{14}t$: $Q_1(1)=\frac{9}{14}$, $Q_1(3)=1-\frac{15}{14}=-\frac{1}{14}$. Check: $\x_1=\x_0-\frac5{14}(1,3)=(\frac9{14},-\frac1{14})$. The error ratio is $\sqrt{(81+3)/196/4}=\sqrt{3/28}\approx0.327$.

After 5 CG steps, $\x_5=\x_0+P(A)\g_0$ and $\e_5=Q(A)\e_0$. What are the largest possible degrees of $Q$ and of $P$?

$\x_5-\x_0\in\mathcal K_4(\g_0)=\operatorname{span}\{\g_0,\dots,A^4\g_0\}$, and $Q(t)=1+tP(t)$.

$\x_5-\x_0$ is a combination of $\d_0,\dots,\d_4$, which span $\mathcal K_4(\g_0)$, so $\deg P\le4$. Then $Q=1+tP$ has degree $\le5$.

$A$ has eigenvalues $1,2,4$. Gradient descent with steps $h_0=\tfrac12$, $h_1=\tfrac14$ has error polynomial $Q(t)=(1-\tfrac t2)(1-\tfrac t4)$. Using this $Q$ in the min–max bound, what upper bound do you get on CG's $\norm{\e_2}_A/\norm{\e_0}_A$?

Evaluate $|Q|$ at each eigenvalue and take the largest.

$Q(1)=\tfrac12\cdot\tfrac34=\tfrac38$, $Q(2)=0$, $Q(4)=(-1)\cdot0=0$. So $\norm{\e_2}_A\le0.375\norm{\e_0}_A$ for CG, from any start. (Gradient descent with these steps gets the same bound for itself; CG can only do better.)

Which of these could be CG's error polynomial $Q_2$ after two steps?

Two conditions: degree at most 2, and value 1 at $t=0$.

$1-t+\frac{t^2}4=(1-\frac t2)^2$ has degree 2 and $Q(0)=1$. $t^2-3t+2$ and $2-t$ have $Q(0)=2$; $(1-t)^3$ has $Q(0)=1$ but degree 3, which needs three steps.

  • Write $\g_0=A\e_0$ (because $\b=A\x^\star$) to pass from $\x_0+P(A)\g_0$ to $Q(A)\e_0$
  • Check both conditions on a residual polynomial: the degree, and $Q(0)=1$
  • Measure the error in the $A$-norm, where the eigen-expansion has no cross terms
  • Remember any admissible $Q$ gives an upper bound; CG's own $Q$ is at least as good
  • Saying CG's polynomial is "monic" or "the Chebyshev polynomial": it is whatever minimizes $\norm{Q(A)\e_0}_A$, and depends on $\e_0$
  • Mixing up $P_k$ (degree $k$, position) and $Q_{k+1}$ (degree $k+1$, error)
  • Forgetting the weights $\lambda_i$ in $\norm{Q(A)\e_0}_A^2=\sum\lambda_iQ(\lambda_i)^2\xi_i^2$
  1. CG's iterates satisfy $\e_k=Q_k(A)\e_0$ with $\deg Q_k\le k$ and $Q_k(0)=1$, and CG's $Q_k$ minimizes $\norm{Q(A)\e_0}_A$ over all such polynomials.
  2. In the eigenbasis, $Q(A)$ multiplies the error component along $\v_i$ by $Q(\lambda_i)$, so $\norm{Q(A)\e_0}_A^2=\sum_i\lambda_iQ(\lambda_i)^2\xi_i^2$.
  3. Min–max bound: $\norm{\e_k}_A\le\max_i|Q(\lambda_i)|\,\norm{\e_0}_A$ for every $Q\in\mathcal Q_k$, so bounding CG means designing a polynomial that is small at the eigenvalues.

After 3 CG steps, $\e_3=Q(A)\e_0$, where $Q$ is…

the Chebyshev polynomial $T_3$
Chebyshev polynomials are a tool for bounding CG. Does CG's polynomial depend on the starting point?
a monic cubic
Which coefficient does the identity $\e=\e_0+P(A)A\e_0$ pin down: the leading one or the constant one?
the polynomial of degree $\le3$ with $Q(0)=1$ that minimizes $\norm{Q(A)\e_0}_A$
Right: this is the optimality theorem, a consequence of the Expanding Subspace Theorem.
a polynomial of degree $\le2$ with $Q(0)=0$
That is closer to the "position" polynomial $tP(t)$. What is $Q$ at $t=0$?

The key identity $\g_0=A\e_0$ holds because…

$\g_0$ is orthogonal to $\e_0$
That is not true in general, and it isn't an identity of the form we need.
$\b=A\x^\star$, so $\g_0=A\x_0-\b=A(\x_0-\x^\star)$
The minimizer solves $A\x^\star=\b$; subtract.
CG chooses $\d_0=-\g_0$
The choice of $\d_0$ is about the algorithm; the identity is about the gradient of a quadratic.

Why does $\norm{\sum_ic_i\v_i}_A^2$ have no cross terms $c_ic_j$ with $i\ne j$?

Because $A$ is positive definite
Positive definiteness makes the terms $\lambda_ic_i^2$ nonnegative. Something else kills the cross terms.
Because the $\xi_i$ are nonnegative
The $\xi_i$ can have any sign. Look at $\v_i^\top A\v_j$.
Because $\v_i^\top A\v_j=\lambda_j\v_i^\top\v_j=0$ for $i\ne j$
Eigenvectors of a symmetric matrix can be chosen orthonormal, and $A\v_j=\lambda_j\v_j$.

You find $Q\in\mathcal Q_4$ with $\max_i|Q(\lambda_i)|=0.3$. What is the sharpest conclusion about CG's $f(\x_4)-f^\star$?

$\le0.3\,(f(\x_0)-f^\star)$
True, but not the sharpest. How is $f-f^\star$ related to $\norm{\e}_A$?
$\le0.09\,(f(\x_0)-f^\star)$
$f-f^\star=\tfrac12\norm{\e}_A^2$, so the factor $0.3$ on $\norm{\e}_A$ squares to $0.09$.
$\le0.3$ in the Euclidean norm $\norm{\e_4}\le0.3\norm{\e_0}$
The bound is for the $A$-norm. Converting to the Euclidean norm can cost a factor $\sqrt\kappa$.

What the spectrum decides: few eigenvalues, clusters, and gradient descent beaten

Because the min–max bound only sees where the eigenvalues sit, the shape of the spectrum, not just $n$ or $\kappa$, decides how fast CG is.

A classic exam question: "$A$ has eigenvalues …; in at most how many steps does CG terminate? Write down the polynomial that proves it." Another: why is CG never worse than gradient descent?

A marksman with $k$ shots: with $k$ targets, every target can be hit. If the targets stand in a few tight groups, one shot per group almost does the job.

If $A$ has only $r$ distinct eigenvalues $\tau_1<\dots<\tau_r$, then CG reaches $\x^\star$ in at most $r$ iterations, from any $\x_0$.

Prove the theorem with an explicit polynomial.

  1. Let $Q(t)=\Big(1-\dfrac t{\tau_1}\Big)\Big(1-\dfrac t{\tau_2}\Big)\cdots\Big(1-\dfrac t{\tau_r}\Big)$.

    One factor per distinct eigenvalue, each normalized so that it equals 1 at $t=0$. All $\tau_j>0$ since $A\succ0$, so the divisions are fine.

  2. $\deg Q=r$ and $Q(0)=1$, so $Q\in\mathcal Q_r$.

    These are the two conditions for an admissible residual polynomial after $r$ steps.

  3. Every eigenvalue $\lambda_i$ equals some $\tau_j$, so $Q(\lambda_i)=0$ for all $i$, and $\max_i|Q(\lambda_i)|=0$.

    Repeated eigenvalues cost nothing: one root kills every copy.

  4. The min–max bound after $r$ steps gives $\norm{\e_r}_A\le0\cdot\norm{\e_0}_A$, so $\x_r=\x^\star$. (If CG reaches $\g_k=\0$ earlier, it has already stopped at $\x^\star$.)

    $\norm{\cdot}_A$ is a norm because $A\succ0$, so $\norm{\e_r}_A=0$ forces $\e_r=\0$.

[NW] writes the same polynomial as $Q_r(\lambda)=\frac{(-1)^r}{\tau_1\cdots\tau_r}(\lambda-\tau_1)\cdots(\lambda-\tau_r)$; expanding the product shows they agree.

Identity plus low rank

Let $A=I+\u\u^\top$ with $\u\ne\0$. Every $\mathbf w\perp\u$ satisfies $A\mathbf w=\mathbf w+\u(\u^\top\mathbf w)=\mathbf w$, so it is an eigenvector with eigenvalue 1; and $A\u=\u+\u\norm{\u}^2$ gives eigenvalue $1+\norm{\u}^2$. Two distinct eigenvalues: CG solves $A\x=\b$ in 2 steps, whatever $n$ is. More generally, $I$ plus a rank-$p$ symmetric positive semidefinite matrix needs at most $p+1$ steps.

This happens in practice. Ridge regression with $p$ data points and $n\gg p$ features (Part 1) solves $(\mu I+X^\top X)\x=X^\top\mathbf y$ with $X$ of size $p\times n$. $X^\top X$ has rank at most $p$, so the matrix has eigenvalue $\mu$ with multiplicity at least $n-p$ and at most $p+1$ distinct eigenvalues: CG needs at most $p+1$ steps. Steepest descent enjoys no such property.

$A$ is $100\times100$ with eigenvalues $1$ (97 times), $10$, $50$ and $200$. At most how many CG steps until $\x_k=\x^\star$?

100
That's the general finite-termination guarantee, but it ignores the repeated eigenvalue.
4
Four distinct eigenvalues, so $Q(t)=(1-t)(1-\frac t{10})(1-\frac t{50})(1-\frac t{200})$ kills them all.
3
The three outliers need three roots, but the eigenvalue 1 needs one too.
Can't tell without $\x_0$
The bound holds for every $\x_0$ (some starting points finish even earlier).
Try it

Here $n=12$ but only three distinct values. Step $k$ from 0 to 3 and watch $Q_k$ (blue, top) acquire roots at the three values; the error (bottom, log scale) drops to machine zero at $k=3$. Then drag one dot slightly away from its group: the error is no longer exactly zero at $k=3$, just very small.

Clusters: Luenberger's bound

Exactly repeated eigenvalues are rare. Tight clusters are common, and a root placed inside a tight cluster makes $Q$ tiny on the whole cluster. The following estimate makes this precise.

For $0\le k\le n-1$, $$\norm{\e_{k+1}}_A\le\frac{\lambda_{n-k}-\lambda_1}{\lambda_{n-k}+\lambda_1}\,\norm{\e_0}_A .$$ ([NW] states the squared version, $\norm{\e_{k+1}}_A^2\le\big(\frac{\lambda_{n-k}-\lambda_1}{\lambda_{n-k}+\lambda_1}\big)^2\norm{\e_0}_A^2$.)

How to read it:

  • $k=0$: the bound is $\frac{\lambda_n-\lambda_1}{\lambda_n+\lambda_1}=\frac{\kappa-1}{\kappa+1}$, the one-step steepest descent rate from Part 5. That makes sense: CG's first step is steepest descent with exact line search.
  • Each further step removes the largest remaining eigenvalue from the effective condition number. The polynomial behind it spends $k$ roots on the $k$ largest eigenvalues and one root at the midpoint of what is left.
  • $m$ outliers plus a cluster. If $m$ eigenvalues are large and the other $n-m$ lie in $[1-\epsilon,1+\epsilon]$, take $k=m$: $\lambda_{n-m}\le1+\epsilon$ and $\lambda_1\ge1-\epsilon$, so after $m+1$ steps $\norm{\e_{m+1}}_A\le\frac{2\epsilon}{2}\norm{\e_0}_A=\epsilon\norm{\e_0}_A$. CG spends $m$ steps on the outliers and one on the whole cluster.

[NW] Fig. 5.4 shows exactly this with five large eigenvalues and the rest in $[0.95,1.05]$: a sharp drop around iteration 6, and another at iteration 7, because the matrix "almost" has six distinct eigenvalues. In general, $r$ tight clusters mean CG approximately solves the problem in about $r$ steps ([NW] p.116): a polynomial with one root per cluster is not zero at the eigenvalues, but tiny.

Go deeper: the polynomial behind Luenberger's bound

[NW] states the theorem without proof. Here is the construction. Let $c=\tfrac12(\lambda_1+\lambda_{n-k})$ and $$Q(t)=\Big(1-\frac tc\Big)\prod_{j=n-k+1}^{n}\Big(1-\frac t{\lambda_j}\Big).$$ It has degree $k+1$ and $Q(0)=1$, and it vanishes at the $k$ largest eigenvalues. For each remaining eigenvalue $\lambda_i\in[\lambda_1,\lambda_{n-k}]$: every product factor $1-\lambda_i/\lambda_j$ lies in $[0,1]$ since $\lambda_i\le\lambda_j$, and the first factor satisfies $|1-\lambda_i/c|=|c-\lambda_i|/c\le\tfrac12(\lambda_{n-k}-\lambda_1)/c=\frac{\lambda_{n-k}-\lambda_1}{\lambda_{n-k}+\lambda_1}$. So $\max_i|Q(\lambda_i)|$ is at most that number; apply the min–max bound.

CG beats every gradient method

Let $\y_0=\x_0$ and $\y_{j+1}=\y_j-h_j\grad f(\y_j)$ for any step sizes $h_0,\dots,h_k$ (constant step, exact line search, Armijo, anything). Then $f(\x_{k+1})\le f(\y_{k+1})$.

Proof. From the story in Chapter 9.1, $\y_{k+1}-\x^\star=Q(A)\e_0$ with $Q(t)=\prod_{j=0}^k(1-h_jt)\in\mathcal Q_{k+1}$. CG's error is the minimum of $\norm{Q(A)\e_0}_A$ over all of $\mathcal Q_{k+1}$, so $\norm{\e_{k+1}}_A\le\norm{\y_{k+1}-\x^\star}_A$, and $f-f^\star=\tfrac12\norm{\cdot}_A^2$. $\blacksquare$

(Step sizes chosen by a line search depend on the iterates, but once the run is over they are just numbers $h_j$, and the identity above holds.) [LD] makes a related point: each CG step is at least as good as one steepest descent step taken from the same point $\x_k$.

Recovering the Part 5 rate. On the whole interval $[m,L]$, the degree-1 polynomial $Q(t)=1-\frac{2t}{m+L}$ has maximum $\frac{L-m}{L+m}=\frac{\kappa-1}{\kappa+1}$, attained at both ends. Using $Q(t)^k$ in the min–max bound shows CG is never slower than $\big(\frac{\kappa-1}{\kappa+1}\big)^k$. That is the polynomial of gradient descent with constant step $h=\frac2{m+L}$: all $k$ roots stacked at the midpoint. Chapter 9.3 spreads them out.

Try it

This is a cluster of nine eigenvalues near 1 plus three outliers (4, 7, 9.5). Step $k$ from 0 to 5: the polynomial puts its first roots near the outliers, then one in the cluster, and the error (blue) collapses around $k=4$ while steepest descent (red) crawls. Then try "3 clusters", "spread" and "uniform, κ = 100", and compare how far the blue curve sits below the gold Chebyshev bound.

$A\in\R^{100\times100}$ is SPD with eigenvalues $1$ (multiplicity 97), $10$, $50$, $200$. (a) In at most how many steps does CG terminate, and which polynomial proves it? (b) What does Luenberger's bound give after $1,2,3,4$ steps? (Book, Exercise "Designing the polynomial".)

  1. (a) Four distinct eigenvalues, so at most 4 steps, by $Q(t)=(1-t)(1-\frac t{10})(1-\frac t{50})(1-\frac t{200})$.

    Theorem on few distinct eigenvalues.

  2. Sorted: $\lambda_1=\dots=\lambda_{97}=1$, $\lambda_{98}=10$, $\lambda_{99}=50$, $\lambda_{100}=200$.

    Luenberger's bound uses $\lambda_{n-k}$, so we need the sorted list.

  3. $k=0$ (1 step): $\frac{200-1}{201}\approx0.990$. $k=1$ (2 steps): $\frac{50-1}{51}\approx0.961$. $k=2$ (3 steps): $\frac{10-1}{11}\approx0.818$.

    Each step drops the largest remaining eigenvalue from the ratio.

  4. $k=3$ (4 steps): $\lambda_{97}=1=\lambda_1$, so the bound is $0$, consistent with (a).

    Compare the Chebyshev bound of Chapter 9.3 with $\kappa=200$: after 4 steps it is $2(0.868)^4\approx1.13$, which says nothing. The spectrum's shape matters far more than $\kappa$.

$A=I+\u\u^\top\in\R^{500\times500}$ with $\norm{\u}^2=3$. In at most how many steps does CG solve $A\x=\b$?

Find the eigenvalues: what does $A$ do to vectors orthogonal to $\u$, and to $\u$ itself?

Vectors $\perp\u$ have eigenvalue 1 (multiplicity 499); $\u$ has eigenvalue $1+3=4$. Two distinct eigenvalues, so at most 2 steps, with $Q(t)=(1-t)(1-\frac t4)$.

Ridge regression with $X\in\R^{20\times1000}$ (20 data points, 1000 features) gives the system $(0.1I+X^\top X)\x=X^\top\mathbf y$. In the worst case, how many CG steps are needed (exact arithmetic)?

What is the rank of $X^\top X$? How many eigenvalues of $0.1I+X^\top X$ can differ from $0.1$?

$\operatorname{rank}X^\top X\le20$, so at least $980$ eigenvalues equal $0.1$, and at most 20 others: at most $21$ distinct eigenvalues, hence at most 21 steps.

$A\in\R^{1000\times1000}$ has three outlying eigenvalues $100$, $300$, $500$, and the other 997 eigenvalues lie in $[0.9,1.1]$, with $0.9$ and $1.1$ both attained. Use Luenberger's bound to bound $\norm{\e_1}_A/\norm{\e_0}_A$ and $\norm{\e_4}_A/\norm{\e_0}_A$.

After $k+1$ steps the bound is $\frac{\lambda_{n-k}-\lambda_1}{\lambda_{n-k}+\lambda_1}$. For 4 steps, $k=3$: which eigenvalue is $\lambda_{n-3}$?

One step ($k=0$): $\frac{500-0.9}{500+0.9}=\frac{499.1}{500.9}\approx0.99641$. Four steps ($k=3$): the three outliers are removed and $\lambda_{n-3}=1.1$, so the bound is $\frac{1.1-0.9}{1.1+0.9}=0.1$.

Steepest descent with exact line search, started from $\x_0$ on a convex quadratic, reaches $f(\y_3)-f^\star=0.2$. CG started from the same $\x_0$ has $f(\x_3)-f^\star$…

Where does $\y_3$ lie, and over which set does $\x_3$ minimize $f$?

$\y_3-\x^\star=Q(A)\e_0$ with $Q(t)=(1-h_0t)(1-h_1t)(1-h_2t)\in\mathcal Q_3$, so $\y_3\in\x_0+\mathcal K_2(\g_0)$. CG's $\x_3$ minimizes $f$ over that set, so $f(\x_3)-f^\star\le0.2$.

  • Count distinct eigenvalues, not eigenvalues with multiplicity
  • Write the proof polynomial as $\prod_j(1-t/\tau_j)$ so that $Q(0)=1$ is visible
  • Sort the eigenvalues before using Luenberger's $\lambda_{n-k}$
  • For "CG vs gradient descent", write gradient descent's error as $\prod_j(1-h_jA)\e_0$
  • Writing $Q(t)=(t-\tau_1)\cdots(t-\tau_r)$, which has $Q(0)\ne1$ in general
  • Judging CG's speed by $\kappa$ alone: two matrices with the same $\kappa$ can need 2 steps or hundreds
  • Expecting exact termination in $r$ steps for clusters: tight clusters give small error, not zero
  1. $r$ distinct eigenvalues ⇒ at most $r$ CG steps, proved by $Q(t)=\prod_{j=1}^r(1-t/\tau_j)$.
  2. Luenberger: $\norm{\e_{k+1}}_A\le\frac{\lambda_{n-k}-\lambda_1}{\lambda_{n-k}+\lambda_1}\norm{\e_0}_A$; with $m$ outliers and a cluster of width $2\epsilon$ around 1, about $m+1$ steps reduce the error to $\epsilon$.
  3. CG is never worse than gradient descent with any step sizes, because every such method's error is $Q(A)\e_0$ for some $Q$ in the same family.

$A$ has eigenvalues $2,2,2,5,5,9$. Which polynomial proves that CG terminates in at most 3 steps?

$(t-2)(t-5)(t-9)$
Check the value at $t=0$.
$(1-\frac t2)^3(1-\frac t5)^2(1-\frac t9)$
Admissible, but what degree is it, and how many steps does that prove?
$(1-\frac t2)(1-\frac t5)(1-\frac t9)$
Degree 3, value 1 at 0, zero at every eigenvalue.
$(1-2t)(1-5t)(1-9t)$
Where are the roots of $1-2t$?

Three $1000\times1000$ SPD matrices all have $\kappa=1000$. On which does CG reach relative error $10^{-6}$ fastest?

Eigenvalues spread uniformly over $[1,1000]$
This is the case where CG behaves most like its $\sqrt\kappa$ worst case.
Five eigenvalues in $[100,1000]$, the rest in $[0.95,1.05]$
Fast (about 6–7 steps), but is there something even better in the list?
Eigenvalue $1$ (500 times) and $1000$ (500 times)
Two distinct eigenvalues: exact solution after 2 steps, whatever $\kappa$ is.

Why is CG's $f(\x_k)$ never larger than that of steepest descent after $k$ steps from the same start?

CG takes longer steps
Step length is not the point. Think about the set each iterate lies in.
Steepest descent's $k$-th iterate lies in $\x_0+\mathcal K_{k-1}(\g_0)$, and CG minimizes $f$ over that set
Its error is $\prod_j(1-h_jA)\e_0$, a member of the family CG optimizes over.
The Chebyshev bound is smaller than $(\frac{\kappa-1}{\kappa+1})^k$
Comparing two upper bounds says nothing about the actual iterates.
CG's directions are orthogonal
CG's directions are $A$-conjugate; its gradients are orthogonal. Neither is the reason.

In [NW] Fig. 5.4 (five large eigenvalues, the rest in $[0.95,1.05]$), why doesn't the error become exactly zero at iteration 6?

Rounding errors
Even in exact arithmetic it would not be zero. Count the distinct eigenvalues.
The cluster contains many distinct values, so a degree-6 polynomial can only be tiny on it, not zero
One root in the cluster makes $|Q|$ small across a narrow interval, but zero only at the root.
Luenberger's bound only holds for $k\ge7$
It holds for every $0\le k\le n-1$.

The $\sqrt\kappa$ bound: Chebyshev polynomials and what they mean for iteration counts

If all you know about $A$ is its smallest and largest eigenvalue, the best polynomial is a rescaled Chebyshev polynomial, and it gives CG's headline rate $2\big(\frac{\sqrt\kappa-1}{\sqrt\kappa+1}\big)^k$.

This is the "$\sqrt\kappa$ versus $\kappa$" comparison quoted in every exam and every numerical library: for $\kappa=10^4$, about 100 times fewer iterations than steepest descent, at essentially the same cost per iteration.

Spacing fence posts: gradient descent with a constant step piles all $k$ roots at one point; Chebyshev spreads them across $[m,L]$, packed more tightly near the ends, so the polynomial stays low everywhere.

Suppose we only know $m=\lambda_1$ and $L=\lambda_n$. Every eigenvalue lies in $[m,L]$, so making $Q$ small on the whole interval is enough: $$\max_i|Q(\lambda_i)|\le\max_{t\in[m,L]}|Q(t)|.$$ Gradient descent's best constant-step polynomial $(1-\frac{2t}{m+L})^k$ gives $\big(\frac{\kappa-1}{\kappa+1}\big)^k$ (Chapter 9.2). Can another degree-$k$ polynomial with $Q(0)=1$ be smaller on $[m,L]$? Yes, and the answer is classical.

$T_0(z)=1$, $T_1(z)=z$, $T_{k+1}(z)=2z\,T_k(z)-T_{k-1}(z)$. So $T_2=2z^2-1$, $T_3=4z^3-3z$, and $T_k$ has degree $k$. Two facts:

  • $T_k(\cos\theta)=\cos k\theta$. Hence $|T_k(z)|\le1$ on $[-1,1]$, and it touches $\pm1$ alternately $k+1$ times.
  • For $|z|\ge1$: $T_k(z)=\tfrac12\big[(z+\sqrt{z^2-1})^k+(z-\sqrt{z^2-1})^k\big]$, so it grows fast outside $[-1,1]$.

(Both formulas satisfy the same recurrence and starting values; for the first, $\cos(k+1)\theta+\cos(k-1)\theta=2\cos\theta\cos k\theta$.)

For all $k\ge0$, $$\norm{\e_k}_A\le2\left(\frac{\sqrt\kappa-1}{\sqrt\kappa+1}\right)^k\norm{\e_0}_A .$$

Where it comes from. The map $z(t)=\frac{L+m-2t}{L-m}$ sends $[m,L]$ onto $[-1,1]$ (reversed) and sends $t=0$ to $z_0=\frac{\kappa+1}{\kappa-1}>1$. Take $$Q(t)=\frac{T_k(z(t))}{T_k(z_0)} .$$ It has degree $k$, and $Q(0)=1$ because we divided by $T_k(z_0)$. On $[m,L]$, $|T_k(z(t))|\le1$, so $|Q|\le1/T_k(z_0)$. Since $z_0>1$ and $T_k$ grows fast outside $[-1,1]$, $T_k(z_0)$ is large and $Q$ is small on the whole interval. A short calculation (below) gives $1/T_k(z_0)\le2q^k$ with $q=\frac{\sqrt\kappa-1}{\sqrt\kappa+1}$. The min–max bound finishes the proof. (If $\kappa=1$, then $A=mI$ and CG finishes in one step.)

Go deeper: the calculation, and why Chebyshev is optimal

With $z_0=\frac{\kappa+1}{\kappa-1}$: $z_0^2-1=\frac{(\kappa+1)^2-(\kappa-1)^2}{(\kappa-1)^2}=\frac{4\kappa}{(\kappa-1)^2}$, so $$z_0+\sqrt{z_0^2-1}=\frac{\kappa+2\sqrt\kappa+1}{\kappa-1}=\frac{(\sqrt\kappa+1)^2}{(\sqrt\kappa-1)(\sqrt\kappa+1)}=\frac{\sqrt\kappa+1}{\sqrt\kappa-1}=\frac1q .$$ By the closed form, dropping the positive second term, $T_k(z_0)\ge\tfrac12q^{-k}$, i.e. $1/T_k(z_0)\le2q^k$. (The second term is $\tfrac12q^k$, which is why the bound is nearly exact for moderate $k$.)

Optimality. Among all $Q$ of degree $\le k$ with $Q(0)=1$, the scaled Chebyshev polynomial has the smallest maximum on $[m,L]$. Sketch: if some $\tilde Q$ had $\max|\tilde Q|<1/T_k(z_0)$, then $Q-\tilde Q$ would alternate in sign at the $k+1$ points where $Q=\pm1/T_k(z_0)$, so it would have $k$ roots inside $[m,L]$, plus a root at $t=0$: $k+1$ roots for a nonzero polynomial of degree $\le k$, impossible. This is not needed for the bound itself.

Try it

Blue is the scaled Chebyshev polynomial, red dashed is gradient descent's $(1-\frac{2t}{m+L})^k$; both equal 1 at $t=0$. Raise the degree and press "Zoom in on the interval": the blue curve wiggles with equal-height ripples inside $[m,L]$ and its ripples shrink far faster than the red curve's maximum. Increase $\kappa$ and see both get worse, red much more. Switch to "$T_k$ on $[-1,1]$" to see the equioscillation.

What $\sqrt\kappa$ means for iteration counts

To guarantee $\norm{\e_k}_A\le\varepsilon\norm{\e_0}_A$:

  • Steepest descent (rate $\frac{\kappa-1}{\kappa+1}$, Part 5): need $k\ge\dfrac{\ln(1/\varepsilon)}{\ln\frac{\kappa+1}{\kappa-1}}\approx\dfrac\kappa2\ln\dfrac1\varepsilon$.
  • CG: need $2q^k\le\varepsilon$, i.e. $k\ge\dfrac{\ln(2/\varepsilon)}{\ln\frac{\sqrt\kappa+1}{\sqrt\kappa-1}}\approx\dfrac{\sqrt\kappa}2\ln\dfrac2\varepsilon$.

The approximations use $\ln\frac{x+1}{x-1}=\ln(1+\frac1x)-\ln(1-\frac1x)\approx\frac2x$ for large $x$. So the iteration count is $O(\kappa\log\frac1\varepsilon)$ for steepest descent and $O(\sqrt\kappa\log\frac1\varepsilon)$ for CG.

$\kappa$SD factor $\frac{\kappa-1}{\kappa+1}$CG factor $\frac{\sqrt\kappa-1}{\sqrt\kappa+1}$Iterations for $\varepsilon=10^{-6}$: SD vs CG
100.8180.51969 vs 23
1000.9800.818691 vs 73
10 0000.99980.98069 078 vs 726
Try it

Move $\kappa$ across six orders of magnitude. On these log–log axes the steepest descent count rises with slope 1 and CG's with slope ½: multiplying $\kappa$ by 100 multiplies CG's work by only 10. Change $\varepsilon$ and notice that it only shifts both lines (the $\log\frac1\varepsilon$ factor).

For $\kappa=100$ and $\varepsilon=10^{-6}$, compute the guaranteed iteration counts for steepest descent and CG (error in the $A$-norm).

  1. SD: factor $\frac{99}{101}\approx0.980198$, and $\ln0.980198\approx-0.020001$. Need $k\ge\frac{\ln10^{-6}}{\ln0.980198}=\frac{-13.8155}{-0.020001}\approx690.75$, so $k=691$.

    Solve $(\frac{\kappa-1}{\kappa+1})^k\le\varepsilon$ by taking logarithms; both logs are negative, so the inequality direction is preserved when dividing.

  2. CG: $\sqrt\kappa=10$, factor $q=\frac9{11}\approx0.81818$, $\ln q\approx-0.20067$. Need $2q^k\le10^{-6}$, i.e. $k\ge\frac{\ln(5\times10^{-7})}{\ln q}=\frac{-14.5087}{-0.20067}\approx72.3$, so $k=73$.

    The factor 2 adds $\ln2$ to the numerator, which barely matters.

  3. Ratio $691/73\approx9.5\approx\sqrt{100}$.

    For large $\kappa$ the ratio of counts tends to $\sqrt\kappa\cdot\frac{\ln(1/\varepsilon)}{\ln(2/\varepsilon)}\approx\sqrt\kappa$.

Careful: these are worst-case bounds in the $A$-norm. (1) Clustered spectra do much better (Chapter 9.2). In the book's lab ($n=1000$, eigenvalues uniform in $[1,1000]$), CG reached relative $A$-norm error $10^{-6}$ in 122 iterations although the bound only guarantees it by 230, while steepest descent was still at $1.8\times10^{-3}$ after 200. (2) To convert to the Euclidean norm, use $\norm{\z}\le\norm{\z}_A/\sqrt m$ and $\norm{\z}_A\le\sqrt L\norm{\z}$: $\ \norm{\e_k}\le\sqrt\kappa\cdot2q^k\norm{\e_0}$. (3) For small $k$ the bound can exceed 1 and say nothing, even when CG has already finished.
Try it

Forty eigenvalues spread uniformly over $[0.1,10]$ ($\kappa=100$). Compare the blue CG curve with the gold Chebyshev bound: CG stays below the bound from the start and then bends downward ever more steeply as it "removes" the extreme eigenvalues one by one (Luenberger's bound at work). Then switch to "3 clusters": same picture for the bounds, completely different for CG.

Using the recurrence, compute $T_3(2)$.

$T_2(z)=2z^2-1$ and $T_3(z)=2zT_2(z)-T_1(z)$.

$T_2(2)=7$, $T_3(2)=2\cdot2\cdot7-2=26$. (Check with $T_3=4z^3-3z$: $32-6=26$.)

Let $\kappa=9$ and $k=3$. Compute the exact maximum of the scaled Chebyshev polynomial on $[m,L]$, namely $1/T_3(z_0)$ with $z_0=\frac{\kappa+1}{\kappa-1}$, and compare it with $2q^3$.

$z_0=\frac{10}8=1.25$ and $T_3(z)=4z^3-3z$.

$T_3(1.25)=4(1.953125)-3.75=4.0625$, so $1/T_3(z_0)=\frac{16}{65}\approx0.2462$. With $q=\frac{3-1}{3+1}=\frac12$, $2q^3=0.25$: the bound $2q^k$ is very close to the exact Chebyshev value.

For $\kappa=10^4$ and $\varepsilon=10^{-6}$, find the smallest $k$ for which the worst-case bound guarantees $\norm{\e_k}_A\le\varepsilon\norm{\e_0}_A$, for CG and for steepest descent.

CG: $q=\frac{99}{101}$, solve $2q^k\le10^{-6}$. SD: factor $\frac{9999}{10001}$, solve $(\cdot)^k\le10^{-6}$.

CG: $\ln q=\ln\frac{99}{101}\approx-0.0200007$, $k\ge\frac{\ln(5\times10^{-7})}{\ln q}\approx\frac{-14.5087}{-0.0200007}\approx725.4$, so $726$. SD: $\ln\frac{9999}{10001}\approx-0.00020000$, $k\ge\frac{-13.8155}{-0.00020000}\approx69077.6$, so $69078$.

For $\kappa=100$, how many CG steps does the Chebyshev bound need to guarantee a Euclidean error reduction $\norm{\e_k}\le10^{-6}\norm{\e_0}$?

$\norm{\e_k}\le\sqrt\kappa\cdot2q^k\norm{\e_0}$ with $q=\frac9{11}$.

Need $20q^k\le10^{-6}$, so $k\ge\frac{\ln(5\times10^{-8})}{\ln(9/11)}=\frac{-16.811}{-0.20067}\approx83.8$, i.e. $k=84$ (versus 73 for the $A$-norm).

  • Quote the rate with its norm: $\norm{\e_k}_A\le2q^k\norm{\e_0}_A$, $q=\frac{\sqrt\kappa-1}{\sqrt\kappa+1}$
  • Turn rates into iteration counts with logarithms: $k\approx\frac{\sqrt\kappa}2\ln\frac2\varepsilon$
  • Use Chebyshev when only $m$ and $L$ are known, and Luenberger or distinct-eigenvalue arguments when you know the spectrum's shape
  • Dropping the factor 2 or the square root ($\kappa$ vs $\sqrt\kappa$) in the CG rate
  • Treating the $\sqrt\kappa$ bound as CG's actual behaviour: it is a worst case over spectra in $[m,L]$
  • Applying the $A$-norm bound to $\norm{\e_k}$ without the $\sqrt\kappa$ conversion factor
  1. The scaled Chebyshev polynomial $T_k(z(t))/T_k(z_0)$ has $Q(0)=1$ and maximum $1/T_k(z_0)\le2q^k$ on $[m,L]$, giving $\norm{\e_k}_A\le2\big(\frac{\sqrt\kappa-1}{\sqrt\kappa+1}\big)^k\norm{\e_0}_A$.
  2. Iterations to accuracy $\varepsilon$: $O(\kappa\log\frac1\varepsilon)$ for steepest descent versus $O(\sqrt\kappa\log\frac1\varepsilon)$ for CG; for $\kappa=10^4$ roughly $69\,000$ versus $730$.
  3. The $\sqrt\kappa$ bound is pessimistic for clustered spectra and is stated in the $A$-norm; converting to the Euclidean norm costs a factor $\sqrt\kappa$.

Gradient descent with the constant step $h=\frac2{m+L}$ corresponds to which residual polynomial after $k$ steps?

$T_k\big(\frac{L+m-2t}{L-m}\big)$
That is the Chebyshev construction (before normalizing), not gradient descent.
$\big(1-\frac{2t}{m+L}\big)^k$
Each step multiplies the error by $I-hA$; all $k$ roots sit at the midpoint $\frac{m+L}2$.
$1-\frac{2kt}{m+L}$
Repeating a step multiplies the factors; it does not add them.

In $Q(t)=T_k(z(t))/T_k(z_0)$, why divide by $T_k(z_0)$?

To make $Q$ monic
Which normalization does a residual polynomial need?
To make $Q(0)=1$, since $z(0)=z_0$
Residual polynomials must equal 1 at $t=0$; dividing also shrinks $Q$ on $[m,L]$ because $T_k(z_0)$ is large.
To keep $|Q|\le1$ on $[-1,1]$
$Q$ is a function of $t\in[m,L]$, and $|T_k|\le1$ there already.

$\kappa$ increases from 100 to 10 000 (fixed $\varepsilon$). CG's worst-case iteration count grows by roughly…

100×
That is steepest descent's growth, proportional to $\kappa$.
10×
The count scales like $\sqrt\kappa$: $\sqrt{10^4}/\sqrt{10^2}=10$ (726 vs 73).
2×
That would be logarithmic growth in $\kappa$. Look at the factor $\frac{\sqrt\kappa}2$.

$A\in\R^{1000\times1000}$ has eigenvalues 1 and 1000 only. After 2 steps the Chebyshev bound is $2q^2\approx1.76$. CG's actual $\norm{\e_2}_A$ is…

about $1.76\norm{\e_0}_A$
That's an upper bound, and CG's error never increases in the $A$-norm. What do two distinct eigenvalues imply?
about $0.94\norm{\e_0}_A$
That's one step of the bound's rate $q$. Use the spectrum's shape instead.
$0$
Two distinct eigenvalues: $Q(t)=(1-t)(1-\frac t{1000})$ vanishes on the spectrum. Worst-case bounds can be useless for special spectra.

Which property of CG makes the polynomial view possible?

The gradients are mutually orthogonal
True of CG, but the polynomial view needs to know which set the iterates live in.
$\operatorname{span}\{\d_0,\dots,\d_k\}=\mathcal K_k(\g_0)=\operatorname{span}\{\g_0,A\g_0,\dots,A^k\g_0\}$
So $\x_{k+1}-\x_0$ is a polynomial in $A$ applied to $\g_0$.
$\alpha_k$ is chosen by exact line search
Exact line search alone gives steepest descent, which has a polynomial too, just not the optimal one.

The Expanding Subspace Theorem says $\x_k$ minimizes $f$ over $\x_0+\operatorname{span}\{\d_0,\dots,\d_{k-1}\}$. Minimizing $f$ there is the same as minimizing…

$\norm{\x-\x^\star}$ (Euclidean)
The Euclidean and $A$-norm minimizers over a subspace differ in general.
$\norm{\x-\x^\star}_A$
$f(\x)-f^\star=\tfrac12\norm{\x-\x^\star}_A^2$.
$\norm{\grad f(\x)}$
That would be a different method (minimizing the residual norm).

Steepest descent with exact line search contracts $\norm{\e}_A$ by at most $\frac{\kappa-1}{\kappa+1}$ per step. Which statement about CG is correct?

CG's first step is faster than steepest descent's first step
From the same point, the first CG step is a steepest descent step.
Luenberger's bound at $k=0$ equals $\frac{\kappa-1}{\kappa+1}$, and later steps can only improve on steepest descent
CG's first step is steepest descent; afterwards CG optimizes over a family containing every gradient method.
CG and steepest descent have the same rate once $k\ge2$
CG's worst case is governed by $\sqrt\kappa$, not $\kappa$.

For $f\in\mathcal S_{\mu,L}^{1,1}$, gradient descent with $h=\frac2{\mu+L}$ contracts $\norm{\x_k-\x^\star}$ by $\frac{L-\mu}{L+\mu}$ per step. On a quadratic with spectrum in $[\mu,L]$, this is the maximum on $[\mu,L]$ of…

$|1-ht|$, the degree-1 residual polynomial of one gradient step
At $t=\mu$ and $t=L$ it equals $\frac{L-\mu}{L+\mu}$, and it is smaller in between.
$|T_1(t)|$
$T_1(t)=t$ is not even 1 at $t=0$.
$1/T_1(z_0)$ with $z_0=\frac{\kappa+1}{\kappa-1}$
Numerically equal! But the question asks which polynomial gradient descent uses. (Indeed, at degree 1, Chebyshev and the best gradient step coincide.)

Why is every vector $\mathbf w\perp\u$ an eigenvector of $I+\u\u^\top$?

Because $I+\u\u^\top$ is diagonal
It usually isn't. Compute $(I+\u\u^\top)\mathbf w$ directly.
$(I+\u\u^\top)\mathbf w=\mathbf w+\u(\u^\top\mathbf w)=\mathbf w$
So the eigenvalue is 1, with an $(n-1)$-dimensional eigenspace: CG needs only 2 steps.
Because $\u\u^\top$ has rank $n-1$
$\u\u^\top$ has rank 1.

Lecture 14 (29 Sep): Newton's method (tangent lines, quadratic models, one-step convergence on quadratics), Nesterov's local quadratic convergence theorem and his example of divergence, how practical codes tame Newton, and the quasi-Newton methods SR1, DFP and BFGS that learn curvature from gradients alone. The part ends with every method of the course racing on the same functions.

You need: Part 0b (gradient, Hessian, Taylor, positive definite matrices), Part 3 (optimality conditions; saddle points), Part 5 (gradient descent and the condition number $\kappa$), the Armijo and Wolfe line searches (Part 6, Lecture 7), and Parts 8–9 (conjugate directions and CG) for the comparisons.

Newton's method: follow the tangent, jump to the bottom of the model

Newton's method replaces the function by its best quadratic approximation at the current point and jumps straight to that approximation's stationary point.

Gradient descent knows which way is downhill but not how far to go, so it zig-zags when $\kappa$ is large. Newton also uses curvature: on a quadratic it finishes in one step, whatever $\kappa$ is.

Parking a car by eye: you don't creep forward a fixed amount at a time; you judge the distance and the curve, then move to where you think the space is, and correct once.

Where it comes from: solving an equation with tangent lines

Forget minimization for a moment. You want a number $t^\star$ with $\phi(t^\star)=0$, for some smooth function $\phi:\R\to\R$. You are at a guess $t_k$. Close to $t_k$, the graph of $\phi$ looks like its tangent line: $$\phi(t_k+\Delta t)\approx\phi(t_k)+\phi'(t_k)\,\Delta t.$$ Pretend this approximation is exact and solve $\phi(t_k)+\phi'(t_k)\Delta t=0$ for the move $\Delta t$. That gives Newton's iteration for a root: $$t_{k+1}=t_k-\frac{\phi(t_k)}{\phi'(t_k)}\qquad(\phi'(t_k)\ne0).$$ Geometrically: draw the tangent at $(t_k,\phi(t_k))$ and walk down it to where it crosses the axis.

Example: $\phi(t)=t^2-2$ (root $\sqrt2$) from $t_0=1$ gives $t_1=1.5$, $t_2=1.416667$, $t_3=1.4142157$, $t_4=1.414213562375$. The errors are $0.41,\ 0.086,\ 0.0025,\ 2.1\times10^{-6},\ 1.6\times10^{-12}$: once close, each step roughly doubles the number of correct digits. Chapter 2 makes this precise.

From roots to minima

A minimizer of a smooth $f$ must satisfy the first-order condition $\grad f(\x)=\0$ (Part 3). So minimizing is "finding a root of the gradient". In one variable, apply the tangent-line rule to $\phi=f'$, whose derivative is $f''$: $$x_{k+1}=x_k-\frac{f'(x_k)}{f''(x_k)}.$$ In $n$ variables the gradient is a map $\R^n\to\R^n$, and its "derivative" (its Jacobian matrix) is the Hessian $\hess f$. Linearizing $\grad f(\x_k+\d)\approx\grad f(\x_k)+\hess f(\x_k)\d$ and setting this to $\0$ gives the multivariable rule.

Given $f\in C^2$ (twice continuously differentiable) and a start $\x_0$, repeat:

  1. solve the linear system $\hess f(\x_k)\,\d_k=-\grad f(\x_k)$ for the Newton direction $\d_k$;
  2. set $\x_{k+1}=\x_k+\d_k$.

In one formula, $\x_{k+1}=\x_k-[\hess f(\x_k)]^{-1}\grad f(\x_k)$. In practice the inverse is never formed: one factorizes $\hess f(\x_k)=LDL^\top$ (or Cholesky), which costs about $\tfrac16n^3$ multiplications ([FR] p.44) and also reveals whether the Hessian is positive definite.

The second derivation: minimize the quadratic model

Taylor's theorem (Part 0b) says that near $\x_k$, $f(\x_k+\d)\approx q_k(\d)$, where $\g_k=\grad f(\x_k)$ and $$q_k(\d):=f(\x_k)+\g_k^\top\d+\tfrac12\d^\top\hess f(\x_k)\,\d .$$ This quadratic model $q_k$ is a bowl when $\hess f(\x_k)\succ0$. Its gradient is $\grad q_k(\d)=\grad f(\x_k)+\hess f(\x_k)\d$, which vanishes exactly at the Newton direction. So Newton jumps to the bottom of the local quadratic model ([FR] (3.1.1)). Gradient descent, by contrast, uses the model $f(\x_k)+\grad f(\x_k)^\top\d+\frac1{2\alpha}\norm{\d}^2$, which pretends the curvature is the same, $1/\alpha$, in every direction.

Picture it: gradient descent stands on a hillside and asks only "which way is steepest?". Newton also feels how the slope is changing under its feet, fits a bowl to what it feels, and leaps to the bottom of that bowl. If the hill really is a bowl, one leap is enough.
Try it

Start with $\sqrt{1+t^2}$ in "root of $f'$" mode and press Step a few times from $t_0=0.9$: each tangent to $f'$ lands much closer to 0. Now set $t_0=1.05$ and watch the tangents overshoot further each time. Switch to "minimize $f$" to see the same iterates as jumps to the bottom of dashed parabolas. Then try $t^3/3-2t$ from a negative start (Newton happily converges to a local maximum) and the cycle function from $t_0=0$.

Newton's method is applied to the strictly convex quadratic $f(\x)=\tfrac12\x^\top A\x-\b^\top\x$ ($A\succ0$) from a start far from the minimizer. How many steps does it need?

About $\kappa$ steps, like gradient descent
That's gradient descent's problem: it ignores curvature. What does Newton's model look like when $f$ is itself quadratic?
At most $n$ steps, like CG
CG needs up to $n$ because it learns curvature one direction at a time. Newton is handed all of it at once.
Exactly one step
The quadratic model of a quadratic is the function itself, so the bottom of the model is the true minimizer.

For $f(\x)=\tfrac12\x^\top A\x-\b^\top\x$ with $A\succ0$: $\grad f(\x_0)=A\x_0-\b$ and $\hess f=A$, so $$\x_1=\x_0-A^{-1}(A\x_0-\b)=A^{-1}\b=\x^\star\quad\text{for every }\x_0 .$$ On the running example (fit the line, $A=\begin{pmatrix}28&12\\12&6\end{pmatrix}$, $\b=(46,20)$, $\kappa\approx46.2$), one Newton step from anywhere lands on $(w,c)=(1.5,\tfrac13)$. Gradient descent with the best fixed step needs hundreds of steps for ten digits.

Why Newton doesn't care about $\kappa$: affine invariance

Change variables $\x=S\y$ with an invertible matrix $S$, and let $g(\y)=f(S\y)$. The chain rule gives $\grad g(\y)=S^\top\grad f(\x)$ and $\hess g(\y)=S^\top\hess f(\x)S$. So the Newton step in $\y$ is $$-\big(S^\top\hess f\,S\big)^{-1}S^\top\grad f=-S^{-1}[\hess f]^{-1}\grad f,$$ which is exactly the Newton step in $\x$, mapped back by $S^{-1}$. Stretching or rotating the coordinates changes nothing: the iterates correspond one-to-one. Gradient descent is not like this, because $\grad g=S^\top\grad f$ points somewhere else; that is why its speed depends on $\kappa$ (Part 5) and Newton's doesn't. Newton is steepest descent measured in the Hessian's own metric, the metric that turns the elliptical level sets into circles.

Is the Newton direction downhill?

A direction $\d$ is a descent direction when $\grad f(\x_k)^\top\d\lt0$. For the Newton direction, $$\grad f(\x_k)^\top\d_k=-\grad f(\x_k)^\top[\hess f(\x_k)]^{-1}\grad f(\x_k).$$ If $\hess f(\x_k)\succ0$ then its inverse is also positive definite (eigenvalues $1/\lambda_i>0$), so this is negative whenever $\grad f(\x_k)\ne\0$: descent. If the Hessian is indefinite, the Newton direction may point uphill or towards a saddle point; Part 3 showed Newton walking into a saddle. Newton is a root-finder for $\grad f=\0$, and maxima and saddles are roots too.

Careful: "Newton's method" in this course means the iteration for minimization, $\x_{k+1}=\x_k-[\hess f]^{-1}\grad f$. It uses second derivatives of $f$. The root-finding version $t_{k+1}=t_k-\phi/\phi'$ uses first derivatives of $\phi$; the two are the same method once $\phi=f'$.

[FR] (3.1.3): take one Newton step for $f(\x)=x_1^4+x_1x_2+(1+x_2)^2$ from $\x_0=(0.75,-1.25)^\top$, and check that $f$ decreased.

  1. Gradient: $\grad f=\big(4x_1^3+x_2,\ x_1+2(1+x_2)\big)$. At $\x_0$: $4(0.421875)-1.25=0.4375$ and $0.75+2(-0.25)=0.25$, so $\g_0=(0.4375,\ 0.25)$.

    Differentiate term by term: $x_1x_2$ contributes $x_2$ to $\partial_1f$ and $x_1$ to $\partial_2f$.

  2. Hessian: $\hess f=\begin{pmatrix}12x_1^2&1\\1&2\end{pmatrix}$, so $\hess f(\x_0)=\begin{pmatrix}6.75&1\\1&2\end{pmatrix}$, with determinant $13.5-1=12.5\gt0$ and positive diagonal: positive definite.

    For a $2\times2$ symmetric matrix, positive diagonal and positive determinant mean both eigenvalues are positive, so the Newton direction will be a descent direction.

  3. Solve $\hess f(\x_0)\d_0=-\g_0$: $\d_0=-\frac1{12.5}\begin{pmatrix}2&-1\\-1&6.75\end{pmatrix}\begin{pmatrix}0.4375\\0.25\end{pmatrix}=-\frac1{12.5}\begin{pmatrix}0.625\\1.25\end{pmatrix}=\begin{pmatrix}-0.05\\-0.1\end{pmatrix}$.

    The $2\times2$ inverse is $\frac1{\det}\begin{pmatrix}d&-b\\-c&a\end{pmatrix}$. For bigger systems you'd factorize instead.

  4. $\x_1=\x_0+\d_0=(0.70,\ -1.35)$. Then $f(\x_0)=-0.55859375$ and $f(\x_1)=0.2401-0.945+0.1225=-0.5824$: lower. The gradient norm drops from $0.504$ to $0.022$.

    These match [FR] Table 3.1.1. The next steps give $\norm{\g}=1.4\times10^{-4}$, then $5.8\times10^{-9}$: the quadratic convergence of Chapter 2.

Newton's root iteration for $\phi(t)=t^2-2$ from $t_0=1$. What is $t_2$? (A fraction or a decimal with 5 places.)

$t_{k+1}=t_k-\frac{t_k^2-2}{2t_k}=\frac{t_k}2+\frac1{t_k}$. First compute $t_1$.

$t_1=\tfrac12+1=\tfrac32$. Then $t_2=\tfrac34+\tfrac23=\tfrac{17}{12}\approx1.41667$. ($\sqrt2=1.41421\ldots$)

Minimize $f(x)=x-\ln x$ on $x\gt0$ with Newton's method from $x_0=0.5$. Derive the iteration and give $x_2$.

$f'(x)=1-1/x$ and $f''(x)=1/x^2$, so $x_{k+1}=x_k-(1-1/x_k)x_k^2$. Simplify.

$x_{k+1}=x_k-x_k^2+x_k=2x_k-x_k^2$. From $0.5$: $x_1=1-0.25=0.75$, $x_2=1.5-0.5625=0.9375$. Note $1-x_{k+1}=(1-x_k)^2$: the error is squared at every step ($0.5\to0.25\to0.0625$).

One Newton step for $f(x,y)=x^4+y^2+xy$ from $(1,0)$. Give the new point.

$\grad f=(4x^3+y,\ 2y+x)=(4,1)$ at $(1,0)$, and $\hess f=\begin{pmatrix}12x^2&1\\1&2\end{pmatrix}=\begin{pmatrix}12&1\\1&2\end{pmatrix}$ there. Solve $\hess f\,\d=-\grad f$.

$\det=23$, so $\d=-\frac1{23}\begin{pmatrix}2&-1\\-1&12\end{pmatrix}\begin{pmatrix}4\\1\end{pmatrix}=-\frac1{23}\begin{pmatrix}7\\8\end{pmatrix}$. New point $(1-\tfrac7{23},\ -\tfrac8{23})=(\tfrac{16}{23},-\tfrac8{23})\approx(0.6957,-0.3478)$.

Newton on $f(x)=x-\ln x$ (defined for $x\gt0$, strictly convex there) is started at $x_0=3$. What happens?

Use $x_{k+1}=2x_k-x_k^2$ from the previous problem.

$x_1=6-9=-3$, which is outside the domain $x\gt0$: $\ln(-3)$ is undefined. Since $1-x_{k+1}=(1-x_k)^2$, the method converges exactly when $|1-x_0|\lt1$, i.e. $0\lt x_0\lt2$. Strict convexity does not save pure Newton from a far start.

  • Derive the Newton step both ways: root of the linearized gradient, and minimizer of the quadratic model
  • Solve $\hess f\,\d=-\grad f$ rather than inverting the Hessian
  • Check that $\hess f(\x_k)\succ0$ before trusting the direction to be downhill
  • On a quadratic, say "one step" and prove it in one line
  • Writing $\x_{k+1}=\x_k-\hess f\,\grad f$ (the Hessian must be inverted, i.e. solved with)
  • Assuming Newton finds minima: it finds stationary points, including maxima and saddles
  • Forgetting that the root form $t-\phi/\phi'$ uses $\phi=f'$, so $\phi'=f''$
  1. Newton's step solves $\hess f(\x_k)\d_k=-\grad f(\x_k)$: it is the root of the linearized gradient and the stationary point of the quadratic model $q_k$.
  2. On a strictly convex quadratic it lands on $\x^\star=A^{-1}\b$ in one step, and it is affine invariant, so the condition number does not slow it down.
  3. The direction is downhill when $\hess f(\x_k)\succ0$; otherwise it can go uphill or to a maximum or saddle.

At $\x_k$, $\hess f(\x_k)=\begin{pmatrix}2&0\\0&-2\end{pmatrix}$ and $\grad f(\x_k)=(0,1)$. The Newton direction $\d_k$ and the slope $\grad f^\top\d_k$ are…

$\d_k=(0,-\tfrac12)$, slope $-\tfrac12$: downhill
Solve $\hess f\,\d=-\grad f$ carefully: the second equation is $-2d_2=-1$.
$\d_k=(0,\tfrac12)$, slope $+\tfrac12$: uphill
$-2d_2=-1$ gives $d_2=\tfrac12$, and $\grad f^\top\d=\tfrac12\gt0$. With an indefinite Hessian the Newton direction can go uphill.
$\d_k=(0,-1)$, the negative gradient
That is the gradient-descent direction; Newton rescales it by the inverse Hessian.

Newton's method and gradient descent are both run on $f(\x)=\tfrac12\x^\top A\x$ and on $g(\y)=f(S\y)$ for an invertible $S$. Which statement is true?

Both methods produce corresponding iterates $\x_k=S\y_k$
Check gradient descent: $\grad g=S^\top\grad f$, which is not $S^{-1}$ times anything natural.
Newton's iterates correspond ($\x_k=S\y_k$); gradient descent's generally do not
The Newton step in $\y$ is $S^{-1}$ times the Newton step in $\x$. That's affine invariance.
Neither, because the Hessians differ
They do differ ($S^\top AS$ vs $A$), but for Newton the difference cancels exactly.

In one dimension, the Newton step for minimization is $x_{k+1}=x_k-f'(x_k)/f''(x_k)$. Which picture describes it?

Walk down the tangent line to $f$ until it meets the $x$-axis
That finds a root of $f$, not a minimizer. Which function's root do we want?
Walk down the tangent line to $f'$ until it meets the axis, i.e. jump to the vertex of the osculating parabola of $f$
The root of $f'$'s tangent and the stationary point of $f$'s quadratic model are the same point.
Move a fixed distance downhill
That's gradient descent with a step size. Newton's step length is set by the curvature $f''$.

Why do practical codes solve $\hess f\,\d=-\grad f$ with an $LDL^\top$ or Cholesky factorization instead of computing $[\hess f]^{-1}$?

Because the inverse doesn't exist when $\hess f\succ0$
A positive definite matrix is always invertible.
Because the factorization gives a different, better direction
Both give exactly the same $\d$ in exact arithmetic.
It's cheaper and more stable, and it reveals whether $\hess f$ is positive definite
About $\frac16n^3$ multiplications ([FR] p.44), and a non-positive pivot flags an indefinite Hessian.

How fast Newton is, why, and where it breaks

Close to a minimizer with a positive definite Hessian, Newton's error is squared at every step, so the number of correct digits doubles; far away it can diverge.

The proof of this (Nesterov's Theorem 1.2.5) is examinable, and so is the example where Newton diverges on a perfectly nice convex function. Practical codes add damping and Hessian fixes for exactly that reason.

A strong swimmer near the shore reaches it in a few strokes, but drop the same swimmer far out in a current and strong strokes in the wrong direction only make things worse.

Measuring speed: rates of convergence

Let $\x_k\to\x^\star$ and write $r_k=\norm{\x_k-\x^\star}$ for the error.

  • Linear: $r_{k+1}\le q\,r_k$ for some fixed $q\in(0,1)$ and all large $k$. Each step gains a fixed number of digits. Gradient descent on a strongly convex function (Part 7) and CG's bound (Part 9) are of this type.
  • Superlinear: $r_{k+1}/r_k\to0$. Each step gains more digits than the last.
  • Quadratic: $r_{k+1}\le c\,r_k^2$ for some $c\gt0$ and all large $k$. If $c\,r_k=10^{-2}$ then $c\,r_{k+1}\le10^{-4}$, then $10^{-8}$, $10^{-16}$: the number of correct digits doubles each step.

Quadratic $\Rightarrow$ superlinear $\Rightarrow$ linear. (Slower still is sublinear, like the $O(1/k)$ bound of Part 7.)

Try it

Compare the columns: for Newton on $t-\ln t$ the "correct decimals" bar doubles and $e_{k+1}/e_k^2$ stays near 1, while gradient descent gains about one digit every three steps with $e_{k+1}/e_k$ stuck near 0.5. Then look at $t^4$, whose minimizer is degenerate: Newton is suddenly only linear.

Nesterov's example: Newton can diverge on a convex function

Find the root $t^\star=0$ of $\phi(t)=\dfrac{t}{\sqrt{1+t^2}}$. (This is $f'$ for the strictly convex $f(t)=\sqrt{1+t^2}$, so it is also Newton for minimizing $f$.) By the quotient rule, $\phi'(t)=\dfrac{\sqrt{1+t^2}-t^2/\sqrt{1+t^2}}{1+t^2}=\dfrac1{(1+t^2)^{3/2}}$, so $$\begin{aligned}t_{k+1}&=t_k-\frac{t_k}{\sqrt{1+t_k^2}}\,(1+t_k^2)^{3/2}\\&=t_k-t_k(1+t_k^2)=-t_k^3 .\end{aligned}$$ If $|t_0|\lt1$, then $|t_k|=|t_0|^{3^k}\to0$ extremely fast (from $0.9$: $0.9,\ -0.729,\ 0.387,\ -0.058,\ 2.0\times10^{-4},\ -7.6\times10^{-12}$). If $|t_0|=1$, it oscillates between $\pm1$ forever. If $|t_0|\gt1$, it diverges (from $1.05$: $-1.16,\ 1.55,\ -3.73,\ 52,\ -1.4\times10^5$).

The lesson: Newton's method is a local method. Its guarantees need the start to be close enough to $\x^\star$. Far out, $\sqrt{1+t^2}$ is nearly linear ($f''$ is tiny), so the quadratic model is a very flat parabola whose bottom lies far beyond the true minimizer.

The local convergence theorem

Nesterov works under three assumptions ([Y] p.34):

  1. (N1) $f\in C^{2,2}_M(\R^n)$: $f$ is twice differentiable and its Hessian is Lipschitz with constant $M$: $\norm{\hess f(\x)-\hess f(\y)}\le M\norm{\x-\y}$ for all $\x,\y$, where $\norm{\cdot}$ on matrices is the operator norm (largest stretch factor; for symmetric matrices, the largest $|\lambda|$). $M$ measures how far $f$ is from quadratic; $M=0$ for a quadratic.
  2. (N2) There is a local minimizer $\x^\star$ with $\hess f(\x^\star)\succeq\ell I$ for some $\ell\gt0$ ([Y] (1.2.22)): the bowl is genuinely curved in every direction at the bottom. ($\ell$ plays the role that $\mu$ played for strong convexity in Part 7, but only at $\x^\star$.)
  3. (N3) The start $\x_0$ is close enough to $\x^\star$.

Under (N1)–(N2), if $\norm{\x_0-\x^\star}\lt\bar r:=\dfrac{2\ell}{3M}$, then $\norm{\x_k-\x^\star}\lt\bar r$ for all $k$, every Newton step is well defined, and $$\norm{\x_{k+1}-\x^\star}\le\frac{M\norm{\x_k-\x^\star}^2}{2\big(\ell-M\norm{\x_k-\x^\star}\big)} .$$ In particular $r_{k+1}\le\frac{3M}{2\ell}r_k^2$: quadratic convergence. ([FR] Thm 3.1.1 states the same fact with less explicit constants.)

Prove Theorem 1.2.5. (Examinable: know steps A–E and why $\bar r=2\ell/(3M)$.) Write $r_k=\norm{\x_k-\x^\star}$ and $H_k=\hess f(\x_k)$.

  1. Step A (exact error formula). Since $\grad f(\x^\star)=\0$, the fundamental theorem of calculus along the segment from $\x^\star$ to $\x_k$ gives $$\grad f(\x_k)=\grad f(\x_k)-\grad f(\x^\star)=\int_0^1\hess f\big(\x^\star+\tau(\x_k-\x^\star)\big)(\x_k-\x^\star)\,d\tau .$$ Therefore $$\x_{k+1}-\x^\star=\x_k-\x^\star-H_k^{-1}\grad f(\x_k)=H_k^{-1}G_k(\x_k-\x^\star),\quad G_k:=\int_0^1\big[H_k-\hess f(\x^\star+\tau(\x_k-\x^\star))\big]d\tau .$$

    Differentiate $\tau\mapsto\grad f(\x^\star+\tau(\x_k-\x^\star))$ by the chain rule to get the integrand. Then write $\x_k-\x^\star=H_k^{-1}H_k(\x_k-\x^\star)$ to pull out the common factor $H_k^{-1}$.

  2. Step B ($G_k$ is small). The two points compared inside $G_k$ are $\x_k$ and $\x^\star+\tau(\x_k-\x^\star)$, which differ by $(1-\tau)(\x_k-\x^\star)$. By (N1), $$\norm{G_k}\le\int_0^1M(1-\tau)r_k\,d\tau=\frac M2r_k .$$

    The norm of an integral is at most the integral of the norm, and $\int_0^1(1-\tau)d\tau=\tfrac12$.

  3. Step C ($H_k$ is safely invertible). For any unit vector $\u$, $\u^\top H_k\u\ge\u^\top\hess f(\x^\star)\u-\norm{H_k-\hess f(\x^\star)}\ge\ell-Mr_k$. So if $r_k\lt\ell/M$, every eigenvalue of $H_k$ is at least $\ell-Mr_k\gt0$, and $\norm{H_k^{-1}}\le\dfrac1{\ell-Mr_k}$.

    This is [Y] Corollary 1.2.2 with (N2). The smallest eigenvalue is the minimum of $\u^\top H\u$ over unit $\u$ (Part 0b), and the inverse has eigenvalues $1/\lambda_i$.

  4. Step D (combine). $r_{k+1}\le\norm{H_k^{-1}}\,\norm{G_k}\,r_k\le\dfrac{Mr_k^2}{2(\ell-Mr_k)}$.

    Take norms in Step A and use $\norm{AB\v}\le\norm A\norm B\norm\v$.

  5. Step E (staying in the ball; where $2\ell/3M$ comes from). Suppose $r_k\lt\bar r=\frac{2\ell}{3M}$. Then $Mr_k\lt\frac23\ell$, so $\ell-Mr_k\gt\frac\ell3$ (in particular $r_k\lt\ell/M$, so Step C applies), and $$r_{k+1}\le\frac{Mr_k^2}{2\ell/3}=r_k\cdot\frac{3Mr_k}{2\ell}\lt r_k\lt\bar r .$$ By induction from $r_0\lt\bar r$, this holds for every $k$, and $r_{k+1}\le\frac{3M}{2\ell}r_k^2$.

    $\bar r$ is exactly the radius at which the contraction factor $\frac{Mr_k}{2(\ell-Mr_k)}$ equals 1: solve $Mr=2(\ell-Mr)$ to get $r=\frac{2\ell}{3M}$.

Careful: the theorem is a sufficient condition. For $\sqrt{1+t^2}$, $\ell=f''(0)=1$ and the Lipschitz constant of $f''$ is $M=\max|f'''|\approx0.859$, so $\bar r\approx0.776$. Newton actually converges for all $|t_0|\lt1$: the guaranteed ball is smaller than the true basin, but it is never wrong.

[Y] (p.36) notes that this ball has almost the same size as the region where gradient descent is guaranteed to converge linearly ($\bar r=2\ell/M$). Hence the standard plan: use a globally safe method to get close, then let Newton finish.

Making Newton practical

  • Damped Newton ([Y] p.34): $\x_{k+1}=\x_k-h_k[\hess f(\x_k)]^{-1}\grad f(\x_k)$, with the step $h_k$ chosen by a line search (Armijo backtracking, Lecture 7) early on and $h_k=1$ near the solution. Whenever $\hess f\succ0$ the direction is downhill, so the line search always finds an acceptable step.
  • Hessian modification ([FR] pp.47–49): if $\hess f(\x_k)$ is not positive definite, replace it by a positive definite matrix, e.g. $\hess f(\x_k)+\nu I$ with $\nu\gt0$ large enough (Levenberg–Marquardt style), or a modified $LDL^\top$ factorization. As $\nu\to\infty$ the step turns into a short gradient-descent step.
  • Cost: each step needs the $\tfrac12n(n+1)$ second derivatives and an $n\times n$ solve (about $\frac16n^3$ multiplications). For $n=10^6$ that is out of the question. This is what quasi-Newton methods (Chapter 3) remove.
  • Degenerate minimizers: if $\hess f(\x^\star)$ is singular, (N2) fails and the rate can drop to linear. For $f(x)=x^4$: $x_{k+1}=x_k-\frac{4x_k^3}{12x_k^2}=\tfrac23x_k$.
Try it

Fit-the-line: pure Newton lands on $(1.5,\tfrac13)$ in one step from anywhere, while GD + Armijo crawls (drag the start: the purple ellipse is Newton's model, and its centre is the next iterate). Rosenbrock: pure Newton is fast from $(-1.2,1)$ but jumps wildly; the damped version follows the valley. Saddle trap: from $(1.5,0.3)$ pure Newton converges to the saddle $(0,0)$ in 3 steps because the Hessian there is indefinite; the modified version goes to a minimizer. $\sqrt{1+x^2}+\sqrt{1+y^2}$ from $(1.3,0.6)$: pure Newton diverges along $x$ (Example 1.2.4 in each coordinate).

For $f(t)=\sqrt{1+t^2}$ (minimizer $t^\star=0$), compute $\bar r=\frac{2\ell}{3M}$ from Theorem 1.2.5. Use $\ell=f''(0)$ and $M=\max_t|f'''(t)|$ (a Lipschitz constant for $f''$ by the mean value theorem).

$f''(t)=(1+t^2)^{-3/2}$ and $f'''(t)=-3t(1+t^2)^{-5/2}$. Maximize $3t(1+t^2)^{-5/2}$ by setting its derivative to zero: you'll get $t^2=\tfrac14$.

$\ell=f''(0)=1$. Differentiating $3t(1+t^2)^{-5/2}$ gives $3(1+t^2)^{-7/2}(1+t^2-5t^2)$, zero at $t=\tfrac12$, so $M=\tfrac32(1.25)^{-5/2}\approx0.8587$. Then $\bar r=\frac2{3(0.8587)}\approx0.776$. The theorem guarantees quadratic convergence for $|t_0|\lt0.776$; the true basin, from $t_{k+1}=-t_k^3$, is $|t_0|\lt1$.

Newton's method on $f(x)=x^4$ from $x_0=3$. Find $x_3$. (Then ask yourself which hypothesis of Theorem 1.2.5 fails.)

$x_{k+1}=x_k-\frac{4x_k^3}{12x_k^2}$.

$x_{k+1}=\tfrac23x_k$, so $x_3=3\cdot\tfrac8{27}=\tfrac89\approx0.889$. The convergence is only linear (ratio $\tfrac23$) because $f''(0)=0$: no $\ell\gt0$ exists, so (N2) fails.

In [Y] Example 1.2.4 with $t_0=0.95$, what is the smallest $k$ with $|t_k|\lt10^{-6}$?

$|t_k|=0.95^{3^k}$. Take logarithms: you need $3^k\,|\ln0.95|\gt\ln10^6$.

$\ln10^6/|\ln0.95|=13.816/0.0513\approx269.3$. Since $3^5=243\lt269.3\lt729=3^6$, the answer is $k=6$ ($|t_5|=0.95^{243}\approx3.9\times10^{-6}$, $|t_6|\approx6\times10^{-17}$). Slow at first (close to 1 the cubing barely helps), then explosive.

Suppose (N1)–(N2) hold with $M=1$, $\ell=1$, and $r_k=0.1$. What upper bound does Theorem 1.2.5 give for $r_{k+1}$? Is $r_k$ inside the ball $\bar r$?

$\bar r=\frac2{3}$. Plug into $\frac{Mr_k^2}{2(\ell-Mr_k)}$.

$0.1\lt\bar r=0.667$, so the theorem applies, and $r_{k+1}\le\frac{0.01}{2(0.9)}=\frac1{180}\approx0.00556$. The next bound is about $\frac{3.1\times10^{-5}}{1.99}\approx1.5\times10^{-5}$: digits doubling.

  • State all three hypotheses of Thm 1.2.5 (Lipschitz Hessian $M$, $\hess f(\x^\star)\succeq\ell I$, $r_0\lt2\ell/3M$) before using it
  • Reproduce the identity $\x_{k+1}-\x^\star=H_k^{-1}G_k(\x_k-\x^\star)$ and the bound $\norm{G_k}\le\frac M2r_k$
  • Use damping and a Hessian fix when starting far away
  • Diagnose the rate from the ratios $r_{k+1}/r_k$ and $r_{k+1}/r_k^2$
  • Calling Newton "globally convergent": Example 1.2.4 diverges on a strictly convex function
  • Expecting quadratic convergence at a degenerate minimizer like $x^4$
  • Reading $\bar r$ as the exact basin: it is only a guaranteed region
  1. Quadratic convergence means $r_{k+1}\le c\,r_k^2$: the number of correct digits doubles each step.
  2. [Y] Thm 1.2.5: with an $M$-Lipschitz Hessian and $\hess f(\x^\star)\succeq\ell I$, Newton converges quadratically from any $\x_0$ with $\norm{\x_0-\x^\star}\lt\frac{2\ell}{3M}$; the proof bounds $\norm{H_k^{-1}}\le\frac1{\ell-Mr_k}$ and $\norm{G_k}\le\frac M2r_k$.
  3. Pure Newton is local: $t_{k+1}=-t_k^3$ diverges for $|t_0|\gt1$. Practical Newton adds a line search and makes the Hessian positive definite.

An algorithm's errors are $10^{-1},\ 10^{-2},\ 10^{-4},\ 10^{-8}$. The convergence looks…

linear with ratio 0.1
Linear would give $10^{-3}$ after $10^{-2}$. Look at what happens to the exponent.
quadratic
Each error is the square of the previous one: $r_{k+1}=r_k^2$.
sublinear
Sublinear is slower than any geometric sequence; this is much faster.

In the proof of Thm 1.2.5, where does the hypothesis $\hess f(\x^\star)\succeq\ell I$ enter?

In bounding $\norm{G_k}\le\frac M2r_k$
That bound only uses the Lipschitz constant $M$.
In showing $\hess f(\x_k)\succeq(\ell-Mr_k)I$, so $\norm{[\hess f(\x_k)]^{-1}}\le\frac1{\ell-Mr_k}$
Without $\ell\gt0$ the Hessian near $\x^\star$ could be nearly singular and its inverse huge.
In the integral formula for $\grad f(\x_k)$
That formula only needs $\grad f(\x^\star)=\0$ and the chain rule.

Pure Newton on $f(x,y)=x^2+\tfrac14y^4-\tfrac12y^2$ from $(1.5,\,0.3)$ converges to $(0,0)$. Why?

Because $(0,0)$ is the global minimizer
$f(0,\pm1)=-\tfrac14\lt0=f(0,0)$. What kind of point is $(0,0)$?
Because the step size was too large
Pure Newton has no step size to tune; the direction itself is the issue.
$(0,0)$ is a nondegenerate stationary point (a saddle), and Newton is attracted to stationary points; at $y=0.3$ the Hessian is indefinite
$\partial_{yy}f=3y^2-1\lt0$ there, so the model is saddle-shaped and the step heads to its stationary point.

Damped Newton with Armijo backtracking is used on a strictly convex $C^2$ function. Compared with pure Newton, it…

is safe from far away and still converges quadratically near $\x^\star$, where the full step $h_k=1$ gets accepted
Each Newton direction is downhill (Hessian PD), the line search forces decrease, and near the solution the unit step passes the Armijo test.
is always slower, because it takes shorter steps
Near the solution it takes exactly the same steps as pure Newton.
only converges linearly
That happens only if the step stays below 1 forever; Armijo starting from $h=1$ accepts the full step when it's good.

Quasi-Newton methods: learning curvature from gradients

Quasi-Newton methods keep a matrix $H_k$ that approximates the inverse Hessian and correct it after every step using only how the gradient changed.

They need no second derivatives, cost $O(n^2)$ per step instead of $O(n^3)$, always go downhill, and still converge superlinearly. BFGS is the default unconstrained optimizer in most software.

Learning a city by walking: every street you walk teaches you how far things really are in that direction. After enough walks in enough directions, your mental map is accurate.

The idea

Every step gives free curvature information. In [FR] notation, with $\g_k=\grad f(\x_k)$, define the step and the gradient change ([FR] (3.2.2)–(3.2.3)): $$\boldsymbol\delta_k=\x_{k+1}-\x_k,\qquad\boldsymbol\gamma_k=\g_{k+1}-\g_k .$$ By Taylor's theorem for the gradient, $\boldsymbol\gamma_k=G\boldsymbol\delta_k+o(\norm{\boldsymbol\delta_k})$, where $G$ is the Hessian ([FR] (3.2.4)). For a quadratic with Hessian $G$ this is exact: $\boldsymbol\gamma_k=G\boldsymbol\delta_k$ ([FR] (3.2.8)). So the true inverse Hessian maps $\boldsymbol\gamma_k$ to $\boldsymbol\delta_k$. In one dimension, $\gamma/\delta$ is just a finite-difference estimate of $f''$: this is the old secant method.

Notation for this chapter ([FR] §3.2)
$H_k$
symmetric matrix approximating $[\hess f]^{-1}$, the inverse Hessian; usually $H_1=I$
$B_k=H_k^{-1}$
the corresponding approximation of the Hessian itself
$\d_k=-H_k\g_k$
the search direction ([FR] calls it $\mathbf s^{(k)}$)
$\boldsymbol\delta_k,\ \boldsymbol\gamma_k$
step $\x_{k+1}-\x_k$ and gradient change $\g_{k+1}-\g_k$
$G$
the Hessian ($[FR]$'s letter); constant for a quadratic

Choose $\x_1$ and a symmetric positive definite $H_1$ (often $I$). For $k=1,2,\dots$:

  1. set $\d_k=-H_k\g_k$;
  2. line search along $\d_k$ for $\alpha_k$, and set $\x_{k+1}=\x_k+\alpha_k\d_k$;
  3. update $H_k$ to $H_{k+1}$ using $\boldsymbol\delta_k$ and $\boldsymbol\gamma_k$.

Advantages over Newton ([FR] p.50): (i) only first derivatives; (ii) $H_k\succ0$ makes $\d_k$ a descent direction, since $\g_k^\top\d_k=-\g_k^\top H_k\g_k\lt0$, even where the true Hessian is indefinite; (iii) $O(n^2)$ multiplications per step, because we update an inverse and never solve a system.

Careful: following [FR] and the course notes, quasi-Newton iterations are counted from 1 ($\x_1$, $H_1$), while Newton's were counted from 0. With $H_1=I$, the first quasi-Newton step is a steepest-descent step.

The quasi-Newton (secant) condition

We cannot ask $H_k$ to satisfy $H_k\boldsymbol\gamma_k=\boldsymbol\delta_k$: the step was taken with $H_k$ before $\boldsymbol\gamma_k$ was known. So we demand it of the next matrix.

$$H_{k+1}\boldsymbol\gamma_k=\boldsymbol\delta_k,$$ or equivalently, for $B=H^{-1}$, $B_{k+1}\boldsymbol\delta_k=\boldsymbol\gamma_k$. This is $n$ equations for the $\tfrac12n(n+1)$ unknowns of a symmetric $H_{k+1}$, so for $n\ge2$ there are many solutions. Each choice of update formula is a different method. In one dimension it fixes $H_{k+1}=\delta_k/\gamma_k$ uniquely: the secant method.

Rank one: SR1

Try the simplest symmetric correction, a rank-one matrix: $H_{k+1}=H+a\u\u^\top$ (dropping the index $k$ on the right, as [FR] does, (3.2.6)). The secant condition requires $$H\boldsymbol\gamma+a\,\u(\u^\top\boldsymbol\gamma)=\boldsymbol\delta,\ \text{ i.e. }$$ $$a(\u^\top\boldsymbol\gamma)\,\u=\boldsymbol\delta-H\boldsymbol\gamma .$$ So $\u$ must be parallel to $\boldsymbol\delta-H\boldsymbol\gamma$. Take $\u=\boldsymbol\delta-H\boldsymbol\gamma$; then $a\,\u^\top\boldsymbol\gamma=1$, which gives the symmetric rank-one formula ([FR] (3.2.7)): $$H_{k+1}^{\rm SR1}=H+\frac{(\boldsymbol\delta-H\boldsymbol\gamma)(\boldsymbol\delta-H\boldsymbol\gamma)^\top}{(\boldsymbol\delta-H\boldsymbol\gamma)^\top\boldsymbol\gamma}.$$

Let $f$ be quadratic with positive definite Hessian $G$. If the SR1 update is well defined at every step and $\boldsymbol\delta_1,\dots,\boldsymbol\delta_n$ are linearly independent, then $H_{n+1}=G^{-1}$, so the method terminates in at most $n+1$ line searches.

Proof. We prove the hereditary property $H_i\boldsymbol\gamma_j=\boldsymbol\delta_j$ for all $j\lt i$ ([FR] (3.2.9)) by induction on $i$. For $i=2$ it is the secant condition. Assume it for $i$, and let $\u_i=\boldsymbol\delta_i-H_i\boldsymbol\gamma_i$. For $j=i$ the secant condition gives $H_{i+1}\boldsymbol\gamma_i=\boldsymbol\delta_i$. For $j\lt i$, $$H_{i+1}\boldsymbol\gamma_j=H_i\boldsymbol\gamma_j+\u_i\frac{\u_i^\top\boldsymbol\gamma_j}{\u_i^\top\boldsymbol\gamma_i},$$ and $$\begin{aligned}\u_i^\top\boldsymbol\gamma_j&=\boldsymbol\delta_i^\top\boldsymbol\gamma_j-\boldsymbol\gamma_i^\top H_i\boldsymbol\gamma_j\\&=\boldsymbol\delta_i^\top G\boldsymbol\delta_j-\boldsymbol\gamma_i^\top\boldsymbol\delta_j\\&=\boldsymbol\delta_i^\top G\boldsymbol\delta_j-\boldsymbol\delta_i^\top G\boldsymbol\delta_j=0,\end{aligned}$$ using the symmetry of $H_i$, the induction hypothesis $H_i\boldsymbol\gamma_j=\boldsymbol\delta_j$, and $\boldsymbol\gamma=G\boldsymbol\delta$ twice. So $H_{i+1}\boldsymbol\gamma_j=H_i\boldsymbol\gamma_j=\boldsymbol\delta_j$. At $i=n+1$: $H_{n+1}G\boldsymbol\delta_j=\boldsymbol\delta_j$ for $n$ independent vectors $\boldsymbol\delta_j$, hence $H_{n+1}G=I$. $\square$

The proof never used exact line searches. But SR1 has two flaws ([FR] p.53): $H_{k+1}$ need not stay positive definite (so $\d_k$ may stop being downhill), and the denominator $(\boldsymbol\delta-H\boldsymbol\gamma)^\top\boldsymbol\gamma$ can be zero or tiny. Practical codes skip the update in that case.

Rank two: DFP

Allow two terms, $H_{k+1}=H+a\u\u^\top+b\v\v^\top$. The secant condition becomes $H\boldsymbol\gamma+a\u(\u^\top\boldsymbol\gamma)+b\v(\v^\top\boldsymbol\gamma)=\boldsymbol\delta$. The natural choice ([FR] (3.2.10)) is one term that produces $\boldsymbol\delta$ and one that cancels $H\boldsymbol\gamma$: $\u=\boldsymbol\delta$, $\v=H\boldsymbol\gamma$, with $a=1/\boldsymbol\delta^\top\boldsymbol\gamma$ and $b=-1/\boldsymbol\gamma^\top H\boldsymbol\gamma$: $$H_{k+1}^{\rm DFP}=H+\frac{\boldsymbol\delta\boldsymbol\delta^\top}{\boldsymbol\delta^\top\boldsymbol\gamma}-\frac{H\boldsymbol\gamma\boldsymbol\gamma^\top H}{\boldsymbol\gamma^\top H\boldsymbol\gamma}$$ ([FR] (3.2.11); Davidon 1959, Fletcher and Powell 1963). Secant check: $H_{k+1}\boldsymbol\gamma=H\boldsymbol\gamma+\boldsymbol\delta\frac{\boldsymbol\delta^\top\boldsymbol\gamma}{\boldsymbol\delta^\top\boldsymbol\gamma}-H\boldsymbol\gamma\frac{\boldsymbol\gamma^\top H\boldsymbol\gamma}{\boldsymbol\gamma^\top H\boldsymbol\gamma}=\boldsymbol\delta$. Read it as: add the new information along $\boldsymbol\delta$, remove what $H$ used to say along $\boldsymbol\gamma$.

If $H\succ0$ and $\boldsymbol\delta^\top\boldsymbol\gamma\gt0$, then $H_{k+1}^{\rm DFP}\succ0$. Hence if $\boldsymbol\delta_k^\top\boldsymbol\gamma_k\gt0$ for all $k$ and $H_1\succ0$, every $H_k$ is positive definite.

Prove [FR] Theorem 3.2.2. (Examinable.)

  1. Since $H\succ0$, write $H=LL^\top$ with $L$ invertible (Cholesky). Take any $\z\ne\0$ and set $\a=L^\top\z$, $\b=L^\top\boldsymbol\gamma$. Then $\z^\top H\z=\a^\top\a$, $\z^\top H\boldsymbol\gamma=\a^\top\b$ and $\boldsymbol\gamma^\top H\boldsymbol\gamma=\b^\top\b$.

    The factorization turns every $H$-weighted product into an ordinary dot product, so Cauchy–Schwarz becomes available.

  2. Therefore $$\z^\top H_{k+1}\z=\underbrace{\a^\top\a-\frac{(\a^\top\b)^2}{\b^\top\b}}_{\ge0}+\underbrace{\frac{(\z^\top\boldsymbol\delta)^2}{\boldsymbol\delta^\top\boldsymbol\gamma}}_{\ge0}.$$

    The first bracket is $\ge0$ by Cauchy–Schwarz, $(\a^\top\b)^2\le\norm\a^2\norm\b^2$. The second is a square divided by the positive number $\boldsymbol\delta^\top\boldsymbol\gamma$.

  3. Both can vanish only if $\a\parallel\b$, i.e. $\z=c\boldsymbol\gamma$ with $c\ne0$ ($L^\top$ is invertible). But then the second term is $c^2(\boldsymbol\gamma^\top\boldsymbol\delta)^2/\boldsymbol\delta^\top\boldsymbol\gamma=c^2\,\boldsymbol\delta^\top\boldsymbol\gamma\gt0$. So $\z^\top H_{k+1}\z\gt0$ for every $\z\ne\0$. Induction on $k$ finishes the proof.

    The two terms are never zero together: this is exactly where $\boldsymbol\delta^\top\boldsymbol\gamma\gt0$ is needed.

  • Quadratic with $G\succ0$: $\boldsymbol\delta^\top\boldsymbol\gamma=\boldsymbol\delta^\top G\boldsymbol\delta\gt0$.
  • Exact line search along a descent direction: $\g_{k+1}^\top\boldsymbol\delta_k=0$ and $\g_k^\top\boldsymbol\delta_k\lt0$, so $\boldsymbol\delta^\top\boldsymbol\gamma=-\g_k^\top\boldsymbol\delta_k\gt0$.
  • Wolfe line search (Part 6, Lecture 7): the curvature condition $\g_{k+1}^\top\d_k\ge c_2\,\g_k^\top\d_k$ with $c_2\lt1$ gives $\boldsymbol\delta^\top\boldsymbol\gamma=\alpha_k(\g_{k+1}-\g_k)^\top\d_k\ge\alpha_k(c_2-1)\g_k^\top\d_k\gt0$.

This is why quasi-Newton methods are paired with Wolfe line searches: they guarantee that DFP and BFGS stay positive definite.

BFGS, and duality

$$\begin{aligned}H_{k+1}^{\rm BFGS}=H&+\Big(1+\frac{\boldsymbol\gamma^\top H\boldsymbol\gamma}{\boldsymbol\delta^\top\boldsymbol\gamma}\Big)\frac{\boldsymbol\delta\boldsymbol\delta^\top}{\boldsymbol\delta^\top\boldsymbol\gamma}\\&-\frac{\boldsymbol\delta\boldsymbol\gamma^\top H+H\boldsymbol\gamma\boldsymbol\delta^\top}{\boldsymbol\delta^\top\boldsymbol\gamma}.\end{aligned}$$ Equivalently, for $B=H^{-1}$: $\displaystyle B_{k+1}^{\rm BFGS}=B+\frac{\boldsymbol\gamma\boldsymbol\gamma^\top}{\boldsymbol\gamma^\top\boldsymbol\delta}-\frac{B\boldsymbol\delta\boldsymbol\delta^\top B}{\boldsymbol\delta^\top B\boldsymbol\delta}$.

Secant check (write $\rho=\boldsymbol\delta^\top\boldsymbol\gamma$, a scalar): $H_{k+1}\boldsymbol\gamma=H\boldsymbol\gamma+\big(1+\frac{\boldsymbol\gamma^\top H\boldsymbol\gamma}{\rho}\big)\boldsymbol\delta-\frac{\boldsymbol\delta(\boldsymbol\gamma^\top H\boldsymbol\gamma)+H\boldsymbol\gamma\,\rho}{\rho}=\boldsymbol\delta$. In $B$-form it's even quicker: $B_{k+1}\boldsymbol\delta=B\boldsymbol\delta+\boldsymbol\gamma-B\boldsymbol\delta=\boldsymbol\gamma$.

Duality. The $B$-form of BFGS is the DFP formula with $H\leftrightarrow B$ and $\boldsymbol\delta\leftrightarrow\boldsymbol\gamma$ swapped ([FR] p.55: the two are dual or complementary; (3.2.14) is the dual $B$-form of DFP). Because Theorem 3.2.2's proof only needs a positive definite matrix and $\boldsymbol\gamma^\top\boldsymbol\delta\gt0$, it applies to $B_{k+1}^{\rm BFGS}$ too, so $H^{\rm BFGS}_{k+1}=(B^{\rm BFGS}_{k+1})^{-1}\succ0$ whenever $\boldsymbol\delta^\top\boldsymbol\gamma\gt0$.

Properties of DFP and BFGS ([FR] p.54, proved in [FR] §3.3–3.4). On a quadratic with exact line searches: (i) termination in at most $n$ iterations with $H_{n+1}=G^{-1}$; (ii) the hereditary property; (iii) the directions are $G$-conjugate, and with $H_1=I$ the iterates are exactly those of conjugate gradients (Parts 8–9). On general functions: (iv) $H_k\succ0$; (v) about $3n^2$ multiplications per iteration; (vi) superlinear local convergence; (vii) global convergence on strictly convex functions with exact line searches. BFGS is the more robust of the two with inexact line searches and is the standard choice.

[FR] Table 3.2.1: $f(\x)=10x_1^2+x_2^2$ (so $G=\mathrm{diag}(20,2)$), $\x_1=(0.1,\,1)^\top$, $H_1=I$, exact line searches. Do the BFGS iteration by hand until it stops.

  1. $\g_1=G\x_1=(2,2)$, $\d_1=-\g_1=(-2,-2)$. Exact step on a quadratic: $\alpha_1=\dfrac{-\g_1^\top\d_1}{\d_1^\top G\d_1}=\dfrac8{80+8}=\dfrac1{11}$. So $\x_2=(0.1-\tfrac2{11},\ 1-\tfrac2{11})=(-\tfrac9{110},\ \tfrac9{11})$ and $\g_2=(-\tfrac{18}{11},\ \tfrac{18}{11})$.

    With $H_1=I$ the first step is steepest descent with exact line search (Part 5's formula).

  2. $\boldsymbol\delta_1=-\tfrac2{11}(1,1)$, $\boldsymbol\gamma_1=G\boldsymbol\delta_1=-\tfrac4{11}(10,1)$, and $\boldsymbol\delta_1^\top\boldsymbol\gamma_1=\tfrac8{121}\cdot11=\tfrac8{11}\gt0$.

    On a quadratic you can get $\boldsymbol\gamma$ as $G\boldsymbol\delta$ instead of subtracting gradients. Always check $\boldsymbol\delta^\top\boldsymbol\gamma\gt0$ before a DFP/BFGS update.

  3. With $H=I$: $\boldsymbol\gamma^\top\boldsymbol\gamma=\tfrac{16}{121}\cdot101=\tfrac{1616}{121}$, so $1+\frac{\boldsymbol\gamma^\top\boldsymbol\gamma}{\boldsymbol\delta^\top\boldsymbol\gamma}=1+\frac{1616/121}{8/11}=\frac{213}{11}$. Then $\frac{\boldsymbol\delta\boldsymbol\delta^\top}{\boldsymbol\delta^\top\boldsymbol\gamma}=\frac1{22}\begin{pmatrix}1&1\\1&1\end{pmatrix}$ and $\frac{\boldsymbol\delta\boldsymbol\gamma^\top+\boldsymbol\gamma\boldsymbol\delta^\top}{\boldsymbol\delta^\top\boldsymbol\gamma}=\frac1{11}\begin{pmatrix}20&11\\11&2\end{pmatrix}$.

    Compute the scalars first, then the three matrices; this keeps the fractions manageable.

  4. $$H_2=I+\frac{213}{242}\begin{pmatrix}1&1\\1&1\end{pmatrix}-\frac1{242}\begin{pmatrix}440&242\\242&44\end{pmatrix}=\frac1{242}\begin{pmatrix}15&-29\\-29&411\end{pmatrix}.$$ Check: $H_2\boldsymbol\gamma_1=\frac1{242}\cdot\big(-\tfrac4{11}\big)(150-29,\ -290+411)=-\tfrac2{11}(1,1)=\boldsymbol\delta_1$. ✓

    Always verify the secant condition: it catches arithmetic slips immediately. (SR1 and DFP give different matrices, $\frac1{382}\begin{pmatrix}21&-19\\-19&381\end{pmatrix}$ and $\frac1{2222}\begin{pmatrix}123&-119\\-119&2301\end{pmatrix}$, but they satisfy the same check.)

  5. $\d_2=-H_2\g_2=\tfrac1{121}(36,-360)$, parallel to $(1,-10)$. Exact step $\alpha_2=\tfrac{11}{40}$ gives $\x_3=(0,0)=\x^\star$. One more update gives $H_3=\mathrm{diag}(\tfrac1{20},\tfrac12)=G^{-1}$.

    $n=2$ iterations, as property (i) promises, and the matrix has learned the inverse Hessian exactly. All three updates produce directions parallel to $(1,-10)$ here, so all three reach $\x^\star$ at step 2.

Try it

Fletcher's example is loaded with BFGS and exact line searches. Slide $k$ from 1 to 3: the purple ellipse of $H_k$ starts as a circle ($H_1=I$), gets the right width along one direction after one update (the secant condition), and coincides with the gold ellipse of $G^{-1}$ at $k=3$. Read off $H_2$ and compare with the worked example. Switch to SR1 and DFP: different $H_2$, same end. Then choose "Armijo (inexact)": termination in $n$ steps is lost, but DFP and BFGS stay positive definite. Try SR1 on a tilted bowl with Armijo steps and watch for a non-positive-definite $H_k$.

$f(\x)=x_1^2+2x_2^2$ (so $G=\mathrm{diag}(2,4)$), $\x_1=(1,1)$, $H_1=I$, exact line search. Compute $H_2$ by the BFGS formula.

$\g_1=(2,4)$, $\alpha_1=\frac{20}{72}=\frac5{18}$, $\boldsymbol\delta_1=-\frac59(1,2)$, $\boldsymbol\gamma_1=-\frac59(2,8)$, $\boldsymbol\delta_1^\top\boldsymbol\gamma_1=\frac{50}9$. The scalar factor is $1+\frac{\boldsymbol\gamma^\top\boldsymbol\gamma}{\boldsymbol\delta^\top\boldsymbol\gamma}=\frac{43}9$.

$\frac{\boldsymbol\delta\boldsymbol\delta^\top}{\boldsymbol\delta^\top\boldsymbol\gamma}=\frac1{18}\begin{pmatrix}1&2\\2&4\end{pmatrix}$ and $\frac{\boldsymbol\delta\boldsymbol\gamma^\top+\boldsymbol\gamma\boldsymbol\delta^\top}{\boldsymbol\delta^\top\boldsymbol\gamma}=\frac1{18}\begin{pmatrix}4&12\\12&32\end{pmatrix}$. So $H_2=I+\frac{43}{162}\begin{pmatrix}1&2\\2&4\end{pmatrix}-\frac9{162}\begin{pmatrix}4&12\\12&32\end{pmatrix}=\frac1{162}\begin{pmatrix}169&-22\\-22&46\end{pmatrix}$. Check: $H_2\boldsymbol\gamma_1=-\frac59\cdot\frac1{162}(338-176,\ -44+368)=-\frac59(1,2)=\boldsymbol\delta_1$ ✓.

$H=I$, $\boldsymbol\delta=(1,0)$, $\boldsymbol\gamma=(\tfrac12,1)$. (These come from the positive definite quadratic with $G=\begin{pmatrix}0.5&1\\1&3\end{pmatrix}$, and $\boldsymbol\delta^\top\boldsymbol\gamma=\tfrac12\gt0$.) Compute the SR1 update and its determinant.

$\u=\boldsymbol\delta-H\boldsymbol\gamma=(\tfrac12,-1)$ and $\u^\top\boldsymbol\gamma=\tfrac14-1=-\tfrac34$.

$H_{\rm new}=I-\tfrac43\begin{pmatrix}\frac14&-\frac12\\-\frac12&1\end{pmatrix}=\begin{pmatrix}\frac23&\frac23\\\frac23&-\frac13\end{pmatrix}$, with determinant $-\tfrac29-\tfrac49=-\tfrac23\lt0$: indefinite, even though $\boldsymbol\delta^\top\boldsymbol\gamma\gt0$. It still satisfies $H_{\rm new}\boldsymbol\gamma=(1,0)=\boldsymbol\delta$. (BFGS on the same pair gives $\begin{pmatrix}6&-2\\-2&1\end{pmatrix}$, determinant 2, positive definite.)

In one dimension, minimize $f(x)=x-\ln x$ by the quasi-Newton iteration $x_{k+1}=x_k-H_kf'(x_k)$ with unit steps. Given $x_0=1.5$ and $x_1=1.2$, compute $H_1$ from the secant condition and then $x_2$.

$f'(x)=1-1/x$: $f'(1.5)=\tfrac13$, $f'(1.2)=\tfrac16$. The secant condition in 1-D: $H_1=\delta/\gamma$.

$\delta=1.2-1.5=-0.3$, $\gamma=\tfrac16-\tfrac13=-\tfrac16$, so $H_1=1.8$ (an estimate of $1/f''$; the true $1/f''(1.2)=1.44$). Then $x_2=1.2-1.8\cdot\tfrac16=0.9$. With $n=1$ the secant condition determines $H$ uniquely, so SR1, DFP and BFGS all give this same secant step.

$H=I$, $\boldsymbol\delta=(2,0)$, $\boldsymbol\gamma=(1,1)$. Which of SR1, DFP, BFGS can be applied?

Compute the SR1 denominator $(\boldsymbol\delta-H\boldsymbol\gamma)^\top\boldsymbol\gamma$, and the DFP/BFGS denominators $\boldsymbol\delta^\top\boldsymbol\gamma$ and $\boldsymbol\gamma^\top H\boldsymbol\gamma$.

$\boldsymbol\delta-\boldsymbol\gamma=(1,-1)$ and $(1,-1)\cdot(1,1)=0$: the SR1 denominator vanishes although the numerator is not zero, so SR1 is undefined (practical codes skip the update). $\boldsymbol\delta^\top\boldsymbol\gamma=2\gt0$ and $\boldsymbol\gamma^\top\boldsymbol\gamma=2\gt0$, so DFP and BFGS are defined and positive definite.

  • Write the secant condition with the new matrix: $H_{k+1}\boldsymbol\gamma_k=\boldsymbol\delta_k$
  • Check $\boldsymbol\delta^\top\boldsymbol\gamma\gt0$ before a DFP/BFGS update, and verify $H_{k+1}\boldsymbol\gamma=\boldsymbol\delta$ after it
  • Use a Wolfe line search with quasi-Newton methods
  • Know which properties need a quadratic and exact line searches (termination, conjugacy) and which hold in general (positive definiteness, $O(n^2)$ cost)
  • Mixing up $H$ (inverse Hessian, $H\boldsymbol\gamma=\boldsymbol\delta$) with $B$ (Hessian, $B\boldsymbol\delta=\boldsymbol\gamma$): DFP in $H$-form looks like BFGS in $B$-form
  • Assuming SR1 keeps $H$ positive definite
  • Claiming $n$-step termination with an inexact line search
  1. Quasi-Newton methods move along $-H_k\g_k$ and update $H_k\approx[\hess f]^{-1}$ so that $H_{k+1}\boldsymbol\gamma_k=\boldsymbol\delta_k$: only gradients, $O(n^2)$ work, superlinear convergence.
  2. SR1 is the unique symmetric rank-one update and has the hereditary property on quadratics, but can lose positive definiteness or break down; DFP and BFGS are dual rank-two updates that stay positive definite whenever $\boldsymbol\delta^\top\boldsymbol\gamma\gt0$, which a Wolfe line search guarantees.
  3. On a quadratic with exact line searches, DFP/BFGS from $H_1=I$ reproduce the CG iterates and end with $H_{n+1}=G^{-1}$.

Why is the secant condition imposed on $H_{k+1}$ and not on $H_k$?

Because $H_k$ is always the identity
Only $H_1$ is usually $I$; later matrices have been updated.
Because $\boldsymbol\gamma_k=\g_{k+1}-\g_k$ is only known after the step taken with $H_k$
The pair $(\boldsymbol\delta_k,\boldsymbol\gamma_k)$ is new information; the next matrix is built to agree with it.
Because $H_k\boldsymbol\gamma_k=\boldsymbol\delta_k$ would make $H_k$ indefinite
It has nothing to do with definiteness: it's a matter of what is known when.

With a Wolfe line search ($c_2\lt1$) along a descent direction $\d_k$, $\boldsymbol\delta_k^\top\boldsymbol\gamma_k$ is…

zero
That's $\g_{k+1}^\top\boldsymbol\delta_k$ under an exact line search, not $\boldsymbol\delta^\top\boldsymbol\gamma$.
possibly negative, so BFGS may fail
The curvature condition rules this out. Write $\boldsymbol\delta^\top\boldsymbol\gamma=\alpha_k(\g_{k+1}-\g_k)^\top\d_k$.
at least $\alpha_k(c_2-1)\g_k^\top\d_k\gt0$
Both $c_2-1$ and $\g_k^\top\d_k$ are negative, so the product is positive and DFP/BFGS stay positive definite.

On a strictly convex quadratic in $\R^n$ with exact line searches and $H_1=I$, BFGS…

needs about $\kappa$ iterations
That's gradient descent's behaviour. BFGS learns the curvature.
produces the same iterates as CG and terminates in at most $n$ iterations with $H_{n+1}=G^{-1}$
[FR] p.54 properties (i)–(iii).
terminates in one iteration like Newton
It starts knowing nothing ($H_1=I$), so its first step is steepest descent.

The $B$-form of BFGS, $B+\frac{\boldsymbol\gamma\boldsymbol\gamma^\top}{\boldsymbol\gamma^\top\boldsymbol\delta}-\frac{B\boldsymbol\delta\boldsymbol\delta^\top B}{\boldsymbol\delta^\top B\boldsymbol\delta}$, is obtained from DFP's $H$-form by…

taking the inverse of each term
Inverses of sums are not sums of inverses. It's a symbolic swap.
transposing
All these matrices are symmetric; transposing changes nothing.
swapping $H\leftrightarrow B$ and $\boldsymbol\delta\leftrightarrow\boldsymbol\gamma$
The two formulas are dual, and the secant condition $H\boldsymbol\gamma=\boldsymbol\delta$ turns into $B\boldsymbol\delta=\boldsymbol\gamma$ under the same swap.

The race: gradient descent, CG, Newton and quasi-Newton compared

Every method in this course trades cost per step against the number of steps; which one wins depends on $n$, on what derivatives you can afford, and on how far you start.

Exam questions often ask you to compare methods or to choose one for a situation, with reasons in terms of cost, storage and rate.

Walking, cycling, driving and flying all get you there; which is fastest depends on the distance, the traffic, and how long it takes to get to the airport.

MethodUsesWork per stepStorageOn a quadraticLocal rate (general $f$)From far away
Gradient descent$\g$$O(n)$ + 1 gradient$O(n)$never exact; factor $\frac{\kappa-1}{\kappa+1}$ per steplinearsafe with Armijo/Wolfe
CG (linear, Parts 8–9)$A\v$ products1 mat-vec + $O(n)$$O(n)$$\le n$ steps; factor $\frac{\sqrt\kappa-1}{\sqrt\kappa+1}$——
Nonlinear CG (Fletcher–Reeves)$\g$$O(n)$ + line search$O(n)$$\le n$ steps (exact l.s.)linear, often fastsafe with strong Wolfe
Newton$\g$, $\hess f$$\frac16n^3$ + Hessian$O(n^2)$1 stepquadraticmay diverge: damp and modify
Quasi-Newton (BFGS)$\g$$\approx3n^2$ + line search$O(n^2)$$\le n$ steps (exact l.s.)superlinearsafe with Wolfe

How the methods are related

  • CG = steepest descent + memory: $\d_k=-\g_k+\beta_k\d_{k-1}$.
  • Newton = steepest descent in the Hessian's metric: it rescales space so the level sets become circles, which is why $\kappa$ disappears.
  • Quasi-Newton = Newton with a learned metric: $H_k$ is assembled from the pairs $(\boldsymbol\delta_k,\boldsymbol\gamma_k)$.
  • Quasi-Newton = CG on quadratics: with $H_1=I$ and exact line searches, DFP and BFGS produce exactly the CG iterates (Lab 14 confirms agreement to about $10^{-16}$ for $n=8$) and also end with $H_{n+1}=G^{-1}$.
  • Quadratic results become local results: near $\x^\star$ every smooth $f$ is close to its quadratic model with $A=\hess f(\x^\star)$, so the quadratic analysis predicts the final phase of every method.
Try it

Rosenbrock from $(-1.2,1)$: GD + Armijo has not converged after 2000 steps, nonlinear CG needs 84, BFGS 35, DFP 37, SR1 43, damped Newton 21 (compare Lab 14: 21, 34, 37, 46). Turn on pure Newton: 6 steps here, but on $\sqrt{1+x^2}+\sqrt{1+y^2}$ it diverges while everything else converges. On Himmelblau, pure Newton converges to the local maximum near $(-0.27,-0.92)$. Finally, type your own $f(x,y)$ (for example (x-1)^2 + 10*(y-x^2)^2) and drag the start around. The lower plot shows $\log_{10}\norm{\grad f}$: straight lines are linear convergence; curves that fall off a cliff are superlinear or quadratic.

You must minimize a smooth function of $n=2000$ variables whose gradient costs about as much as 5 evaluations of $f$, and whose Hessian is available but costs as much as $n$ gradients. Compare one step of Newton and one step of BFGS, and recommend a method.

  1. Newton: the Hessian costs about $2000$ gradient-equivalents, and the factorization about $\frac16n^3\approx1.3\times10^9$ multiplications.

    The $\frac16n^3$ figure is for an $LDL^\top$ factorization ([FR] p.44).

  2. BFGS: one gradient (plus a few more inside the Wolfe line search) and about $3n^2=1.2\times10^7$ multiplications for the update and the product $H_k\g_k$.

    [FR] p.54, property (v). That is roughly 100 times less arithmetic than Newton's solve, before even counting the Hessian.

  3. Newton needs fewer iterations (quadratic vs superlinear), but typically only a few times fewer; Lab 14's Rosenbrock run had 21 vs 34. That cannot repay a factor of about 100 in arithmetic plus 2000 gradients' worth of Hessian. Recommend BFGS with a Wolfe line search.

    Newton wins when the Hessian is cheap and $n$ is small (tens to hundreds), or when very high accuracy is needed quickly. For huge $n$, even $O(n^2)$ storage is too much: use nonlinear CG or limited-memory BFGS.

Go deeper: limited-memory BFGS and superlinear convergence

L-BFGS ([NW] Ch. 9) never stores $H_k$. It keeps only the last $m$ pairs $(\boldsymbol\delta_i,\boldsymbol\gamma_i)$ (typically $m=5$–$20$) and computes $H_k\g_k$ by a two-loop recursion that applies the BFGS update $m$ times to a scaled identity. Work and storage are $O(mn)$, which is why it is the default for large-scale smooth problems, including many machine-learning models.

Why BFGS is superlinear. It is not true that $H_k\to[\hess f(\x^\star)]^{-1}$ in general. The Dennis–Moré condition shows that superlinear convergence only needs $H_k$ to be accurate along the search directions: $\norm{(H_k^{-1}-\hess f(\x^\star))\d_k}/\norm{\d_k}\to0$. The secant condition enforces exactly this kind of accuracy along the steps actually taken ([NW] §6.4). This is beyond the lecture.

Using [FR]'s counts, Newton's factorization costs about $\frac16n^3$ multiplications and a DFP/BFGS update about $3n^2$. For $n=1000$, how many times more arithmetic is the Newton solve than the quasi-Newton update?

The ratio is $\frac{n^3/6}{3n^2}=\frac n{18}$.

$\frac{1000}{18}\approx55.6$. And Newton also needs the $\tfrac12n(n+1)\approx5\times10^5$ second derivatives.

A quadratic has $\kappa=100$. Using the worst-case factors $\frac{\kappa-1}{\kappa+1}$ per step for steepest descent with exact line search and $2\big(\frac{\sqrt\kappa-1}{\sqrt\kappa+1}\big)^k$ for CG, find the smallest $k$ that guarantees the error has shrunk by a factor $10^{6}$ for each method. (Newton needs 1 step; BFGS with exact line searches at most $n$.)

Steepest descent: $(99/101)^k\le10^{-6}$. CG: $2(9/11)^k\le10^{-6}$. Take logarithms.

Steepest descent: $k\ge\frac{\ln10^6}{\ln(101/99)}=\frac{13.82}{0.0200}\approx690.8$, so $k=691$. CG: $k\ge\frac{\ln(2\times10^6)}{\ln(11/9)}=\frac{14.51}{0.2007}\approx72.3$, so $k=73$. The $\sqrt\kappa$ makes a factor-of-ten difference.

Solve $A\x=\b$ where $A$ is a $10^6\times10^6$ sparse symmetric positive definite matrix with about 10 nonzeros per row. Which method?

Can you store an $n\times n$ dense matrix for $n=10^6$? Which method needs only products $A\v$?

Conjugate gradients: each step costs one sparse product (about $10^7$ operations) and $O(n)$ storage, and it converges at the $\sqrt\kappa$ rate. Newton and BFGS need $O(n^2)=10^{12}$ numbers of storage; gradient descent's rate depends on $\kappa$ instead of $\sqrt\kappa$.

$n=20$, the Hessian is available in closed form and cheap, the start may be far from the solution, and you need 14 correct digits. Which method?

Which method combines global safety with the fastest local rate when second derivatives are cheap?

Damped Newton (with a Hessian modification if needed): the line search makes it safe from far away, and near $\x^\star$ it takes full steps and converges quadratically, so the last ten digits take only a few steps. With $n=20$ the $\frac16n^3\approx1300$ multiplications are negligible.

  • Compare methods on four axes: information needed, work per step, storage, rate (local and from far away)
  • Count derivative evaluations, not just iterations
  • Recommend CG for large sparse quadratics, BFGS for moderate $n$ with gradients, damped Newton when Hessians are cheap
  • Judging a method only by its iteration count
  • Recommending Newton or BFGS when $n^2$ numbers cannot even be stored
  • Forgetting that every fast local rate needs a globalization (line search) to be trusted from a far start
  1. Per step: GD and CG cost $O(n)$, BFGS $O(n^2)$, Newton $O(n^3)$ plus a Hessian; locally: linear, linear, superlinear, quadratic.
  2. Newton and BFGS are blind (or nearly blind) to $\kappa$; GD pays $\kappa$ and CG pays $\sqrt\kappa$.
  3. On quadratics with exact line searches, BFGS/DFP from $H_1=I$ coincide with CG, which ties Parts 8–10 together.

On Rosenbrock's function, gradient descent with a line search needs thousands of iterations while BFGS needs about 35. The main reason is…

BFGS uses exact second derivatives
BFGS uses only gradients.
the valley is badly conditioned, and BFGS learns the curvature while GD's rate is governed by $\kappa$
GD's error ratio is close to 1 in the valley (Lab 14 measured 0.999 per step); BFGS's metric adapts to the valley's shape.
GD's line search is inexact
Even with exact line searches GD zig-zags along a narrow valley.

Which method has the lowest storage requirement?

BFGS
It stores an $n\times n$ matrix $H_k$.
Newton
It stores (and factorizes) the $n\times n$ Hessian.
Nonlinear CG
Just a few vectors: $\x_k$, $\g_k$, $\d_{k-1}$: $O(n)$.

Pure Newton diverges on $\sqrt{1+x^2}+\sqrt{1+y^2}$ from $(1.3,0.6)$, but BFGS converges. Why can BFGS succeed where Newton fails?

BFGS uses the exact Hessian more cleverly
BFGS never computes the Hessian.
Its direction is always downhill ($H_k\succ0$) and its line search forces $f$ to decrease, whereas pure Newton accepts the full step blindly
Globalization by a line search is what saves it; damped Newton converges here too.
BFGS converges quadratically
BFGS is superlinear, not quadratic; and the issue here is global behaviour, not local rate.

Newton's method on $f(\x)=\tfrac12\x^\top A\x-\b^\top\x$, $A\succ0$, from $\x_0$. Then $\x_1=$

$\x_0-\alpha(A\x_0-\b)$ for the exact step $\alpha$
That's one step of steepest descent with exact line search.
$A^{-1}\b$
$\x_0-A^{-1}(A\x_0-\b)=A^{-1}\b=\x^\star$, from any start.
$\x_0-A(A\x_0-\b)$
The Hessian has to be inverted, not multiplied.

Pure Newton converges to a point $\bar\x$ with $\grad f(\bar\x)=\0$ and $\hess f(\bar\x)$ having eigenvalues $3$ and $-1$. Then $\bar\x$ is…

a local minimizer, because Newton converged there
Newton converges to any nondegenerate stationary point. Check the second-order condition.
a saddle point
An indefinite Hessian at a stationary point means a saddle (Part 3's second-order test).
a local maximizer
A maximizer needs a negative semidefinite Hessian; one eigenvalue here is positive.

In [Y] Thm 1.2.5, the radius $\bar r=\frac{2\ell}{3M}$ is chosen so that…

the Hessian is invertible, which needs exactly $r\lt\frac{2\ell}{3M}$
Invertibility only needs $r\lt\ell/M$; the theorem asks for more.
the contraction factor $\frac{Mr_k}{2(\ell-Mr_k)}$ is below 1, so the iterates stay in the ball and the error shrinks
$\frac{Mr}{2(\ell-Mr)}\lt1\iff r\lt\frac{2\ell}{3M}$.
the method converges linearly with ratio $\frac23$
Inside the ball the convergence is quadratic.

On a quadratic with $\kappa=10^4$, which method's iteration count is unaffected by $\kappa$?

Gradient descent with exact line search
Its factor is $\frac{\kappa-1}{\kappa+1}$, extremely close to 1 here.
CG
CG's bound depends on $\sqrt\kappa$ (though it also terminates in $n$ steps).
Newton
Affine invariance: one step regardless of conditioning.

BFGS with $H_1=I$ and exact line searches on a strictly convex quadratic generates directions that are…

orthogonal: $\d_i^\top\d_j=0$
The relevant notion is orthogonality after weighting by $G$.
$G$-conjugate, and the same as those of conjugate gradients
[FR] p.54 property (iii); Lab 14 confirms the iterates agree with CG's to rounding error.
all equal to $-\g_k$
Only the first direction is the negative gradient.

Which update can produce an indefinite $H_{k+1}$ from a positive definite $H_k$ even when $\boldsymbol\delta^\top\boldsymbol\gamma\gt0$?

BFGS
BFGS keeps $H\succ0$ whenever $\boldsymbol\delta^\top\boldsymbol\gamma\gt0$ (by duality with Thm 3.2.2).
DFP
That's exactly [FR] Theorem 3.2.2: DFP preserves positive definiteness.
SR1
Its correction $\u\u^\top/\u^\top\boldsymbol\gamma$ has the sign of $\u^\top\boldsymbol\gamma$, which can be negative; example: $H=I$, $\boldsymbol\delta=(1,0)$, $\boldsymbol\gamma=(\tfrac12,1)$.

The whole course (Lectures 1–14) on a few screens: every key statement with a link back to where it's taught, a method chooser, the three proof templates behind every convergence theorem, and a final mixed review. Use it the week before an exam, after you've worked through the parts.

You need: Parts 1–10. If a line below feels unfamiliar, click its link and redo that chapter's practice.

The course on one page, and which method to use when

Fourteen lectures reduce to about thirty statements; knowing exactly what each one assumes is most of the exam.

Exam questions often hand you a function and ask which theorem applies, or which hypothesis fails. That needs the statements and their assumptions side by side.

A pilot's pre-flight checklist: you know how to fly, but under pressure you still read it line by line.

Existence and optimality (Lectures 2–5)

StatementAssumptionsWhere
sup/inf exist; may not be attainednonempty, bounded set of reals2.1
Weierstrass: $f$ attains its min and max on $K$$f$ continuous, $K$ nonempty, closed and bounded2.4
Coercive ($f\to\infty$ as $\norm{\x}\to\infty$) $\Rightarrow$ a global minimizer exists$f$ continuous on $\R^n$2.4
FONC: local min $\Rightarrow\grad f(\x^\star)=\0$$f\in C^1$, interior point3.1
SONC: local min $\Rightarrow\grad f=\0$, $\hess f(\x^\star)\succeq0$$f\in C^2$3.2
SOSC: $\grad f=\0$, $\hess f(\x^\star)\succ0\Rightarrow$ strict local min$f\in C^2$3.2
Indefinite Hessian at a critical point $\Rightarrow$ saddle; singular $\Rightarrow$ the test is silent$f\in C^2$3.3

Convexity (Interlude)

StatementAssumptionsWhere
Chord $\iff$ tangent below: $f(\y)\ge f(\x)+\grad f(\x)^\top(\y-\x)$ $\iff$ $\hess f\succeq0$$C^1$ for the 2nd, $C^2$ for the 3rd; convex domain4.2
Convex: every local min is global; $\grad f(\x^\star)=\0\iff$ global min$f$ convex ($C^1$ for the second)4.3
Strongly convex ($\mu>0$) $\Rightarrow$ a unique minimizer exists$f\in C^1$4.3
$L$-smooth: $\norm{\grad f(\x)-\grad f(\y)}\le L\norm{\x-\y}$; for $C^2$: all Hessian eigenvalues in $[-L,L]$ (for convex $f$: $\hess f\preceq LI$)—4.4
Quadratic sandwich: $\frac\mu2\norm{\y-\x}^2\le f(\y)-f(\x)-\grad f(\x)^\top(\y-\x)\le\frac L2\norm{\y-\x}^2$$f\in\mathcal S^{1,1}_{\mu,L}$4.4

Gradient methods and line search (Lectures 6–10)

StatementAssumptionsWhere
Fixed step on a quadratic converges $\iff0<\alpha<2/L$; best $\alpha=\frac2{m+L}$, error factor $\frac{\kappa-1}{\kappa+1}$ per step$f=\frac12\x^\top A\x-\b^\top\x$, $A\succ0$5.2
Exact step $\alpha_k=\frac{\g_k^\top\g_k}{\g_k^\top A\g_k}$; consecutive gradients orthogonal; $f$-gap factor $\le\big(\frac{\kappa-1}{\kappa+1}\big)^2$ (Kantorovich)same quadratic5.3
Descent lemma: $f(\y)\le f(\x)+\grad f(\x)^\top(\y-\x)+\frac L2\norm{\y-\x}^2$$L$-smooth (no convexity)6.2
Armijo: $f(\x+\alpha\d)\le f(\x)+c_1\alpha\,\g^\top\d$; Wolfe adds $\grad f(\x+\alpha\d)^\top\d\ge c_2\,\g^\top\d$$0\lt c_1\lt c_2\lt1$, $\g^\top\d\lt0$6.3
Backtracking stops, with $\alpha_k\ge\min\{\bar\alpha,\ 2\beta(1-c_1)|\g^\top\d|/(L\norm{\d}^2)\}$$L$-smooth, descent direction6.4
Zoutendijk: $\sum_k\cos^2\theta_k\norm{\g_k}^2\lt\infty$, so $\cos\theta_k\ge\delta\gt0\Rightarrow\norm{\g_k}\to0$bounded below, $L$-smooth, Wolfe steps7.2
Constant step $h$: $f(\x_k)-f(\x_{k+1})\ge h(1-\frac{hL}2)\norm{\g_k}^2$; best $h=1/L$; $\min_k\norm{\g_k}=O(1/\sqrt N)$$L$-smooth, bounded below7.2
[Y] Thm 2.1.14: $f(\x_k)-f^\star\le\frac{2L\norm{\x_0-\x^\star}^2}{k+4}$ at $h=1/L$$f\in\mathcal F^{1,1}_L$ (convex, $L$-smooth)7.3
[Y] Thm 2.1.15: $\norm{\x_k-\x^\star}\le\big(\frac{Q_f-1}{Q_f+1}\big)^k\norm{\x_0-\x^\star}$ at $h=\frac2{\mu+L}$$f\in\mathcal S^{1,1}_{\mu,L}$, $Q_f=L/\mu$7.4

Conjugate gradients, Newton, quasi-Newton (Lectures 11–14)

StatementAssumptionsWhere
Conjugate directions ($\d_i^\top A\d_j=0$) with exact steps reach $\x^\star$ in $\le n$ steps; $\g_k\perp\d_i$ for $i\lt k$ (expanding subspace)$A\succ0$ quadratic8.2
CG: $\alpha_k=\frac{\norm{\g_k}^2}{\d_k^\top A\d_k}$, $\beta_{k+1}=\frac{\norm{\g_{k+1}}^2}{\norm{\g_k}^2}$, $\d_{k+1}=-\g_{k+1}+\beta_{k+1}\d_k$; gradients orthogonal, directions conjugate$A\succ0$ quadratic8.3
$\norm{\x_k-\x^\star}_A\le\min_{Q(0)=1}\max_i|Q(\lambda_i)|\,\norm{\x_0-\x^\star}_A$; $r$ distinct eigenvalues $\Rightarrow\le r$ steps; $2\big(\frac{\sqrt\kappa-1}{\sqrt\kappa+1}\big)^k$$A\succ0$ quadratic9.2, 9.3
Newton: $\x_{k+1}=\x_k-[\hess f(\x_k)]^{-1}\grad f(\x_k)$; one step on a PD quadratic; local quadratic convergenceLipschitz Hessian, $\hess f(\x^\star)\succeq\ell I$, start close ([Y] Thm 1.2.5)10.2
Secant condition $H_{k+1}\boldsymbol\gamma_k=\boldsymbol\delta_k$; SR1, DFP, BFGS updates; PD kept if $\boldsymbol\delta_k^\top\boldsymbol\gamma_k\gt0$[FR] notation10.3
Try it

Describe your problem and get a recommended method, with the reason and the theorem that backs it. Change one answer at a time and see what flips the recommendation.

An exam question: "$f(x)=x^4$. Does gradient descent with constant step $h=1/L$ converge linearly?" Answer it using the sheet.

  1. Linear rates (Thm 2.1.15) need $f\in\mathcal S^{1,1}_{\mu,L}$: strongly convex and $L$-smooth.

    Always start by naming the theorem and listing its hypotheses.

  2. $f''(x)=12x^2$ is unbounded, so $f$ is not $L$-smooth on $\R$: no single $L$ exists, and "$h=1/L$" isn't even defined.

    $L$-smooth needs $|f''|\le L$ everywhere.

  3. Also $f''(0)=0$, so $f$ is not strongly convex near its minimizer.

    Strong convexity needs $f''\ge\mu\gt0$ everywhere.

  4. So the theorem doesn't apply. On a bounded sublevel set $f$ is $L$-smooth and convex, so the $O(1/k)$ result applies there, but not the linear one; in fact near 0 the iterates creep: $x_{k+1}=x_k(1-4hx_k^2)$.

    Saying which hypothesis fails, and what you can still conclude, is the full-marks answer.

Gradient descent with a fixed step on $f=\tfrac12\x^\top A\x$, where $A$ has eigenvalues 1 and 10. Give the stability limit for $\alpha$, the best fixed step, and the error factor per step at that step.

$0\lt\alpha\lt2/L$; best $\alpha=2/(m+L)$; factor $(\kappa-1)/(\kappa+1)$.

$2/L=0.2$; $\alpha^\star=2/11\approx0.182$; factor $(10-1)/(10+1)=9/11\approx0.818$.

$A$ is $1000\times1000$, symmetric positive definite, with only three distinct eigenvalues. At most how many CG iterations (exact arithmetic) solve $A\x=\b$?

Build a polynomial with $Q(0)=1$ that vanishes at each distinct eigenvalue.

$Q(t)=\prod_{i=1}^3(1-t/\lambda_i)$ has degree 3, $Q(0)=1$, and is zero on the whole spectrum, so the error is zero after at most 3 steps.

In a BFGS step, $\boldsymbol\delta_k^\top\boldsymbol\gamma_k=0.3\gt0$ and $H_k\succ0$. What does theory guarantee about $H_{k+1}$?

What condition did DFP/BFGS need to stay positive definite?

With $\boldsymbol\delta^\top\boldsymbol\gamma\gt0$ and $H_k\succ0$, the update keeps $H_{k+1}\succ0$ (and it satisfies the secant condition by construction). It need not equal the true inverse Hessian after one step.

  • Quote a theorem together with its hypotheses
  • When a theorem doesn't apply, say which hypothesis fails and what still holds
  • Translate between "rate factor per step" and "iterations for accuracy ε"
  • Using $\hess f\preceq LI$ as the definition of $L$-smooth for non-convex $f$
  • Applying quadratic-only results (exact step formula, $n$-step termination) to general $f$
  • Mixing up rates for $\norm{\x_k-\x^\star}$ with rates for $f(\x_k)-f^\star$ (the latter is squared)
  1. Every result in the course is "these hypotheses ⇒ this conclusion"; exams test both halves.
  2. Gradient methods: rates depend on $\kappa$; CG: on $\sqrt\kappa$ and the spectrum; Newton: quadratic but only locally.
  3. Pick the method by structure: quadratic → CG; cheap Hessian and small $n$ → Newton; otherwise quasi-Newton or gradient methods with line search.

Which pair is the $f$-gap rate per step for exact-line-search gradient descent on a quadratic, and the distance rate for the best fixed step?

Both $\frac{\kappa-1}{\kappa+1}$
One of them measures $f(\x_k)-f^\star$, which behaves like a squared distance.
$\big(\frac{\kappa-1}{\kappa+1}\big)^2$ and $\frac{\kappa-1}{\kappa+1}$
Function gaps go like squared distances, so their factor is squared.
$\frac{\sqrt\kappa-1}{\sqrt\kappa+1}$ and $\frac{\kappa-1}{\kappa+1}$
The $\sqrt\kappa$ factor belongs to CG, not gradient descent.

Zoutendijk's theorem alone gives $\norm{\g_k}\to0$ when…

any line search is used
The theorem assumes a specific kind of step and a condition on the directions.
Wolfe steps are used and $\cos\theta_k$ stays above some $\delta\gt0$
Then $\delta^2\sum\norm{\g_k}^2\lt\infty$, so the gradients go to 0.
$f$ is convex
Convexity is not one of its hypotheses; it's about step quality and angles.

For a huge sparse SPD linear system, which method is the natural first choice?

Newton's method
For a quadratic, Newton is "solve $A\x=\b$ directly", which is what we can't afford here.
Conjugate gradients
Only matrix–vector products, $O(n)$ memory, and a $\sqrt\kappa$ rate.
Gradient descent with exact line search
It works, but CG is never worse and usually far faster.

Three proof templates behind every convergence theorem

Almost every convergence proof in Lectures 8–10 is one of three short arguments: sum the decreases, invert a self-improving recursion, or iterate a contraction.

If you recognise the template, an unseen convergence proof on the exam becomes a fill-in-the-blanks exercise.

Chess openings: thousands of games, but a handful of opening patterns. Learn the patterns and the moves follow.

Every proof starts the same way: a one-step inequality from the descent lemma or a co-coercivity inequality. What happens next decides the rate.

From $f(\x_k)-f(\x_{k+1})\ge c\,\norm{\g_k}^2$, add the inequalities for $k=0,\dots,N$. The left side telescopes: $$c\sum_{k=0}^N\norm{\g_k}^2\le f(\x_0)-f(\x_{N+1})\le f(\x_0)-f^\star.$$ So $\sum\norm{\g_k}^2\lt\infty$, hence $\norm{\g_k}\to0$, and $\min_{k\le N}\norm{\g_k}^2\le\frac{f(\x_0)-f^\star}{c\,(N+1)}$. With $h=1/L$, $c=\frac1{2L}$. Used for: constant steps ([Y] 1.2.3), backtracking, Zoutendijk.

If $\Delta_k=f(\x_k)-f^\star\gt0$ satisfies $\Delta_{k+1}\le\Delta_k-c\,\Delta_k^2$, then $\Delta_k\le\dfrac{\Delta_0}{1+c\,\Delta_0\,k}$. Used for: [Y] Thm 2.1.14, where convexity turns "decrease $\propto\norm{\g}^2$" into "decrease $\propto\Delta^2$".

If $r_{k+1}^2\le\rho\,r_k^2$ with $\rho\lt1$, then $r_k^2\le\rho^kr_0^2$, and $r_k\le\varepsilon$ after $k\ge\frac{\ln(r_0^2/\varepsilon^2)}{\ln(1/\rho)}$ steps. Used for: [Y] Thm 2.1.15, gradient descent on quadratics, and, with $\sqrt\kappa$, CG.

Prove Template 2: if $\Delta_k\gt0$ and $\Delta_{k+1}\le\Delta_k-c\Delta_k^2$, then $\Delta_k\le\Delta_0/(1+c\Delta_0k)$.

  1. The recursion shows $\Delta_{k+1}\le\Delta_k$: the sequence decreases.

    We'll need $\Delta_k/\Delta_{k+1}\ge1$ in a moment.

  2. Divide $\Delta_{k+1}\le\Delta_k-c\Delta_k^2$ by the positive number $\Delta_k\Delta_{k+1}$: $\ \dfrac1{\Delta_k}\le\dfrac1{\Delta_{k+1}}-c\,\dfrac{\Delta_k}{\Delta_{k+1}}$.

    The trick is to look at $1/\Delta_k$, which turns a quadratic recursion into an additive one.

  3. Rearrange and use $\Delta_k/\Delta_{k+1}\ge1$: $\ \dfrac1{\Delta_{k+1}}\ge\dfrac1{\Delta_k}+c\,\dfrac{\Delta_k}{\Delta_{k+1}}\ge\dfrac1{\Delta_k}+c$.

    Each step adds at least $c$ to the reciprocal.

  4. Add over $k$ steps: $\dfrac1{\Delta_k}\ge\dfrac1{\Delta_0}+ck$, so $\Delta_k\le\dfrac{1}{1/\Delta_0+ck}=\dfrac{\Delta_0}{1+c\Delta_0k}$. $\blacksquare$

    A telescoping sum again, this time on reciprocals. The result is $O(1/k)$: sublinear.

$\Delta_{k+1}\le\Delta_k-\tfrac12\Delta_k^2$ with $\Delta_0=1$. What bound does Template 2 give for $\Delta_{10}$?

$\Delta_k\le\Delta_0/(1+c\Delta_0k)$ with $c=\tfrac12$.

$1/(1+\tfrac12\cdot1\cdot10)=1/6\approx0.167$.

Gradient descent with $h=1/L$ on an $L$-smooth function with $L=4$ and $f(\x_0)-f^\star=8$. After $N+1=100$ iterations, what does Template 1 guarantee for $\min_k\norm{\g_k}^2$?

With $h=1/L$ the decrease constant is $c=1/(2L)$, and $\min\norm{\g}^2\le\frac{f(\x_0)-f^\star}{c(N+1)}$.

$c=1/8$, so the bound is $8/(\tfrac18\cdot100)=64/100=0.64$.

Gradient descent on a strongly convex function with $Q_f=9$ and $h=2/(\mu+L)$. How many iterations guarantee $\norm{\x_k-\x^\star}\le10^{-6}\norm{\x_0-\x^\star}$? (Smallest integer.)

The factor is $(Q_f-1)/(Q_f+1)=0.8$ per step. Solve $0.8^k\le10^{-6}$.

$k\ge\frac{6\ln10}{\ln1.25}=\frac{13.8155}{0.22314}\approx61.9$, so $k=62$.

Newton's method has error $e_0=10^{-2}$ and satisfies $e_{k+1}\le e_k^2$. After how many steps is $e_k\le10^{-16}$ guaranteed?

Square repeatedly: the exponent doubles each step.

$10^{-2}\to10^{-4}\to10^{-8}\to10^{-16}$: 3 steps. Compare Template 3, where a fixed factor $\rho$ would need about $16/\log_{10}(1/\rho)$ steps.

  • Write the one-step inequality first, then decide which template finishes the proof
  • Keep track of whether you're bounding $\norm{\g}$, $f-f^\star$ or $\norm{\x-\x^\star}$
  • Check the sign and positivity conditions before dividing
  • Telescoping without using $f\ge f^\star$ (boundedness below is essential in Template 1)
  • Dividing by $\Delta_{k+1}$ without noting it's positive
  • Claiming a linear rate from Template 1 (it only gives $\norm{\g_k}\to0$, at a slow rate)
  1. Template 1 (sum decreases) ⇒ gradients go to zero; needs $f$ bounded below.
  2. Template 2 ($\Delta_{k+1}\le\Delta_k-c\Delta_k^2$, look at $1/\Delta_k$) ⇒ $O(1/k)$ for convex smooth $f$.
  3. Template 3 (contraction $\rho$) ⇒ linear rate, $\ln(1/\varepsilon)/\ln(1/\rho)$ iterations.

Which template proves [Y] Thm 2.1.15 (strongly convex, linear rate)?

Template 1
Summing decreases gives only $\norm{\g_k}\to0$, no geometric rate.
Template 2
That gives $O(1/k)$, which is slower than linear.
Template 3
Strong co-coercivity gives $r_{k+1}^2\le\rho r_k^2$, iterated to $\rho^kr_0^2$.

In Template 2's proof, the key move is…

studying $1/\Delta_k$, which increases by at least $c$ each step
That turns a quadratic recursion into a simple additive one.
taking logarithms of $\Delta_k$
Logarithms suit geometric (Template 3) decay, not $\Delta^2$ recursions.
bounding $\Delta_k$ by $\Delta_0$
True but too weak; it gives no rate at all.

Template 1 needs which assumption beyond the one-step decrease?

Convexity
Template 1 works for non-convex functions too (it proves $\norm{\g_k}\to0$).
$f$ bounded below, so the telescoped sum is finite
Without $f^\star\gt-\infty$, the sum of decreases could be infinite.
Strong convexity
That's what upgrades to Template 3; Template 1 doesn't need it.

Which condition guarantees that a continuous $f:\R^n\to\R$ has a global minimizer?

$f$ is bounded below
$e^{x}$ is bounded below by 0 and never reaches it.
$f$ is coercive
Coercive ⇒ some sublevel set is nonempty, closed and bounded ⇒ Weierstrass.
$f$ is differentiable
$f(x)=x$ is differentiable and has no minimizer.
$f$ is convex
$e^x$ is convex and has no minimizer.

At a critical point, $\hess f=\begin{pmatrix}2&0\\0&0\end{pmatrix}$. The second-order test says…

strict local minimum
SOSC needs the Hessian positive definite; here one eigenvalue is 0.
saddle point
A saddle needs a negative eigenvalue somewhere.
nothing conclusive: SONC holds but SOSC fails
Higher-order terms decide: compare $x^2+y^4$ (min) with $x^2-y^4$ (saddle).

$f$ is convex and $C^1$, and $\grad f(\x^\star)=\0$. Then $\x^\star$ is…

a local minimizer only
Convexity upgrades this: what does the tangent-plane inequality give at $\x^\star$?
a global minimizer
$f(\y)\ge f(\x^\star)+\0^\top(\y-\x^\star)=f(\x^\star)$ for all $\y$.
the unique minimizer
Uniqueness needs strict convexity: think of a flat-bottomed valley.

Gradient descent with exact line search on a quadratic. Consecutive gradients $\g_k$ and $\g_{k+1}$ are…

orthogonal
At the exact minimizer along $\d_k=-\g_k$, the new gradient has zero slope along $\g_k$: $\g_{k+1}^\top\g_k=0$.
parallel
If they were parallel, the method would converge along a line; it zig-zags instead.
A-conjugate
That's the property CG's directions have, not gradient descent's gradients.

The Wolfe curvature condition exists to rule out…

steps that are too long
Armijo (sufficient decrease) already stops steps from being too long.
steps that are too short
It requires the slope to have flattened enough, so tiny steps fail it.
non-descent directions
The direction is fixed before the line search; the conditions judge the step length.

For convex, $L$-smooth (not strongly convex) $f$, gradient descent with $h=1/L$ has $f(\x_k)-f^\star$…

decreasing geometrically
That needs strong convexity.
bounded by $O(1/k)$
[Y] Thm 2.1.14: $\le2L\norm{\x_0-\x^\star}^2/(k+4)$.
reaching zero in finitely many steps
Finite termination is a CG-on-quadratics phenomenon.

Conjugate direction methods on an $n$-dimensional SPD quadratic finish in at most…

$\kappa$ steps
Finite termination depends on the dimension, not the condition number.
$n$ steps
$n$ conjugate directions span $\R^n$, and each exact step removes one error component.
$\sqrt\kappa$ steps
$\sqrt\kappa$ governs the rate bound, not exact termination.

To reduce the error by a fixed factor with $\kappa=10^4$, CG's bound needs roughly how many times fewer iterations than gradient descent's?

10,000
The comparison is $\kappa$ against $\sqrt\kappa$.
100
$\kappa/\sqrt\kappa=\sqrt\kappa=100$.
2
The gap grows with $\kappa$; at $10^4$ it's large.

Newton's method on $f=\tfrac12\x^\top A\x-\b^\top\x$ with $A\succ0$, from any start, reaches $\x^\star$ in…

one step
The quadratic model is exact, so its minimizer $A^{-1}\b$ is reached immediately.
$n$ steps
That's CG's guarantee.
it depends on $\kappa$
Newton is affine invariant: conditioning doesn't matter here.

The quasi-Newton (secant) condition in [FR] notation is…

$H_{k+1}\boldsymbol\delta_k=\boldsymbol\gamma_k$
$H$ approximates the inverse Hessian, so it maps gradient changes to steps.
$H_{k+1}\boldsymbol\gamma_k=\boldsymbol\delta_k$
$\boldsymbol\gamma=\g_{k+1}-\g_k\approx G\boldsymbol\delta$, so $H\approx G^{-1}$ should map $\boldsymbol\gamma$ to $\boldsymbol\delta$.
$H_{k+1}=H_k+\boldsymbol\delta_k\boldsymbol\gamma_k^\top$
That's a guess at an update formula, not the condition it must satisfy.

An algorithm stops at a point with $\grad f=\0$ on a non-convex $f$. You can conclude…

it's the global minimizer
Without convexity, local information can't certify global optimality.
it's a local minimizer
It could be a saddle or a maximum: check the Hessian.
only that it's a critical point; second-order checks are needed
FONC is necessary, not sufficient.

Ill-conditioning ($\kappa\gg1$) means the level sets of a quadratic are…

nearly circular
That's $\kappa\approx1$.
long, thin ellipses (axis ratio $\sqrt\kappa$)
Half-axes scale like $1/\sqrt{\lambda_i}$.
hyperbolas
Hyperbolas come from indefinite matrices.

Practice built from this course's past papers (midterms and finals from 2015, 2017, 2021 and 2024). The problems are adapted with new numbers, but the style, difficulty and traps match the real exams. If you're preparing for a midsem, start here and use the links to revise whatever trips you up.

You need: Parts 1–10 for full coverage. Each set says which lectures it draws on, so you can practise as soon as you've done those parts.

How the exams are set, and how to answer them

The exams are short, computational and precise: small concrete functions, exact answers (often fractions), and true/false traps on the hypotheses of theorems.

Knowing the format lets you practise the right skill: applying a theorem to numbers quickly and justifying it in two or three lines.

A driving test: you know how to drive, but passing also means knowing exactly what the examiner checks at each junction.

A midterm is 70–90 minutes, closed book, with 4–6 questions of 5–10 marks each, and answers written in boxes. The final is 180 minutes. Across the papers the same question types come back again and again:

Question typeTypical wordingRevise
Classify critical points of a 1-D polynomial"Find all points with $f'(x)=0$; local or global, min or max? Is $f$ coercive?"3.1, 3.2, 2.4
2-D critical points and global minimum"Is $x_1^4+x_2^4-4x_1x_2$ coercive? Find its global minimum."2.4, 3.2
Quadratic given partial information"$Q\v=2\v$, $\norm{\v}=10$, $\kappa=10$: find $\x^\star$, $f^\star$, iterations of exact-line-search descent from $\0$."5.3
Rate as a fraction"Find $\rho$ with $f(\x_{k+1})-f^\star\le\rho\,(f(\x_k)-f^\star)$. Answer as a fraction."5.3
Iteration counting from a decrease bound"Decrease $\ge0.2\norm{\g_k}^2$, $f\ge0$, $f(\x_0)=20$: how many iterations guarantee $\norm{\g}\le0.1$?"7.2
Step-size facts"Gradients of successive iterates are orthogonal. Which step-size rule was used?"5.3, 6.3
CG iteration count"$\b$ is a combination of $k$ eigenvectors; how many CG iterations from $\0$?"9.2
Quasi-Newton reasoning"Rank-1 matrices $G^{(2)}$, $G^{(4)}$ had negative eigenvalues. Is the code buggy?"10.3
True/false on hypotheses"If $f$ is coercive then every minimum is global. T/F"everywhere
Not yet in your course: the 2015 and 2017 second midterms and finals, and the 2021 third midterm, are mostly about constrained optimization: KKT conditions, Lagrange duality, projections onto boxes and balls, gradient projection, active-set methods and the simplex method. In the 2024 course these came after quasi-Newton methods (October–November). They aren't in Lectures 1–14, so they're not practised here yet.

How to write answers that get full marks

  • Name the result and check its hypotheses in one line: "$f$ is a convex quadratic with $Q\succ0$, so…".
  • Give exact values. If the box says "answer as a fraction", $4/9$ gets marks and $0.44$ may not.
  • Use the symmetric part. $\x^\top M\x=\x^\top\tfrac12(M+M^\top)\x$, so a non-symmetric matrix in a quadratic form is a classic trap.
  • Look for eigenvectors. If the gradient at the start is an eigenvector of $Q$, exact-line-search descent finishes in one step.
  • For true/false, a single counterexample settles "false": keep $x^3$, $x^4$, $-x^4$, $e^x$ and $x^2-y^2$ in your head.
Try it

Practise under real conditions. Start a 70-minute timer, then work through the three practice sets below without hints first. Pause if you need to; the timer remembers nothing once you leave, by design.

Adapted from Midterm 1 (2024), Q2. $f(\x)=\tfrac12\sum_{i=1}^d x_i^2/a_i-\sum_{i=1}^dx_i$ with all $a_i\ne0$, $\sum_ia_i=12$ and $a_{\max}/a_{\min}=3$. Find $f(\hat\x)$ where $\grad f(\hat\x)=\0$, and, if $f$ is bounded below, the $\rho$ in $f(\x_{k+1})-f(\hat\x)\le\rho\,(f(\x_k)-f(\hat\x))$ for steepest descent with exact line search.

  1. $\partial f/\partial x_i=x_i/a_i-1=0$ gives $\hat x_i=a_i$, so $\hat\x=\mathbf a$.

    The function separates into one-variable pieces; each is a parabola.

  2. $f(\hat\x)=\tfrac12\sum a_i-\sum a_i=-\tfrac12\sum a_i=-6$.

    Substitute $x_i=a_i$: $\tfrac12a_i^2/a_i=\tfrac12a_i$.

  3. Bounded below forces every $1/a_i\ge0$ (a negative one would let $f\to-\infty$ along that axis), so all $a_i\gt0$ and the Hessian $\mathrm{diag}(1/a_i)\succ0$.

    This is the hidden step examiners look for: "bounded below" is what makes the quadratic convex.

  4. $\kappa=\frac{1/a_{\min}}{1/a_{\max}}=\frac{a_{\max}}{a_{\min}}=3$, so $\rho=\big(\frac{\kappa-1}{\kappa+1}\big)^2=\big(\tfrac24\big)^2=\tfrac14$.

    Exact-line-search rate in function value (Kantorovich), Part 5.3.

Same setup as the worked example, but now $\sum_ia_i=8$ and $a_{\max}/a_{\min}=7$, and $f$ is bounded below. Find $f(\hat\x)$ and $\rho$ (as a fraction).

$f(\hat\x)=-\tfrac12\sum a_i$, and $\kappa=a_{\max}/a_{\min}$.

$f(\hat\x)=-4$. $\kappa=7$, so $\rho=\big(\tfrac{6}{8}\big)^2=\tfrac{9}{16}$.

  • Write the theorem's name and hypotheses before the computation
  • Keep fractions exact until the last line
  • Symmetrize a matrix before reading off eigenvalues of a quadratic form
  • Forgetting that "bounded below" for a quadratic means the Hessian is positive semidefinite
  • Answering "global minimum" for a coercive function's first critical point without comparing values
  • Spending time on constrained-optimization papers before those lectures happen
  1. Midterms test fast, exact application of theorems to small concrete functions.
  2. Recurring moves: symmetric part, eigenvector start, $\kappa\to\rho$ fractions, decrease bound → iteration count.
  3. True/false items are about hypotheses; carry a pocket set of counterexamples.

An exam gives $f(\x)=\x^\top M\x$ with $M=\begin{pmatrix}1&-2\\2&2\end{pmatrix}$. The Hessian of $f$ is…

$2M$
That's only right when $M$ is symmetric. Is it?
$M+M^\top=\begin{pmatrix}2&0\\0&4\end{pmatrix}$
$\x^\top M\x=\x^\top\tfrac{M+M^\top}2\x$; the off-diagonal parts cancel.
$M$
Differentiate $\x^\top M\x$ twice: a factor 2 appears, and only the symmetric part survives.

"$f$ is a quadratic and bounded below." What can you conclude about its Hessian $Q$?

$Q\succ0$
Too strong: $x_1^2$ in two variables is bounded below with a singular Hessian.
$Q\succeq0$
A negative eigenvalue would let $f\to-\infty$ along its eigenvector.
Nothing
Try $f=-x^2$: it isn't bounded below. What does that say about the sign of $Q$?

Which topic from the past papers is not part of Lectures 1–14?

Counting CG iterations from the eigen-structure of $\b$
That's Lecture 13 (Part 9).
KKT conditions and Lagrange duality
Those are constrained-optimization topics, which come later in the course.
Rank-1 quasi-Newton updates
That's Lecture 14 (Part 10).

Set A: critical points, coercivity, convexity

The first midterm always opens with "find and classify the critical points" and "is it coercive?", usually with one twist.

These are the easiest marks on the paper, if the classification is airtight.

A warm-up lap: everyone runs it, but the people who pace it well have more left for later.

Try each problem without the hint first. Lectures used: 2–5 and the convexity interlude (Parts 2–4).

(Midterm 1, 2024 style.) $f(x)=\tfrac34x^4-x^3+2$. Find the global minimizer and the minimum value.

$f'(x)=3x^2(x-1)$. Check $f''$ at each critical point, and whether $f$ is coercive.

Critical points $x=0,1$. $f''(x)=9x^2-6x$: $f''(1)=3\gt0$ (local min), $f''(0)=0$ (test silent). $f$ is coercive (leading term $\tfrac34x^4$), so a global minimum exists and is at a critical point: $f(1)=\tfrac34-1+2=1.75$, $f(0)=2$. Global minimizer $x=1$, value $1.75$.

For the same $f(x)=\tfrac34x^4-x^3+2$, what is $x=0$?

$f''(0)=0$, so look at the sign of $f'$ just left and right of 0.

$f'(x)=3x^2(x-1)\lt0$ on both sides of 0 (near 0), so $f$ is decreasing through 0: neither a minimum nor a maximum. Equivalently, $f(x)-f(0)=x^3(\tfrac34x-1)$ changes sign at 0.

(Midterm 1, 2017 style.) $f(x)=x^2-\tfrac23x^3$. Find the value of $f$ at its local maximum.

$f'(x)=2x(1-x)$; $f''(x)=2-4x$.

Critical points 0 and 1; $f''(0)=2\gt0$ (local min, value 0), $f''(1)=-2\lt0$ (local max), $f(1)=1-\tfrac23=\tfrac13$.

For $f(x)=x^2-\tfrac23x^3$ on $\R$, what about a global minimum? (Also ask yourself: is $\{x: f(x)\ge M\}$ empty for large $M$? Is $f$ coercive?)

What happens to $-\tfrac23x^3$ as $x\to+\infty$?

$f(x)\to-\infty$ as $x\to+\infty$, so there's no global minimum (and $f$ isn't coercive). The set $\{f\ge M\}$ is never empty: $f\to+\infty$ as $x\to-\infty$.

(Midterm 1, 2015 style.) $f(x_1,x_2)=x_1^4+x_2^4-8x_1x_2$. It is coercive. Find its global minimum value.

Critical points: $4x_1^3=8x_2$ and $4x_2^3=8x_1$. Substitute $x_2=x_1^3/2$.

$x_2=x_1^3/2$ and $x_1^9/8=2x_1$ give $x_1=0$ or $x_1^8=16$, i.e. $x_1=\pm\sqrt2$, $x_2=\pm\sqrt2$ (same sign). $f(\pm\sqrt2,\pm\sqrt2)=4+4-16=-8$, $f(0,0)=0$. Since $f$ is coercive, the global minimum is at a critical point: $-8$.

For $f=x_1^4+x_2^4-8x_1x_2$, classify the critical point $(0,0)$.

$\hess f(\0)=\begin{pmatrix}0&-8\\-8&0\end{pmatrix}$. Its determinant?

$\det=-64\lt0$: eigenvalues $\pm8$, indefinite, so $(0,0)$ is a saddle.

(Midterm 1, 2017 style.) $f(\x)=\tfrac12\x^\top A\x-2\b^\top\x$ with $A=\begin{pmatrix}2&1\\-1&a\end{pmatrix}$, $\b=(1,0)^\top$. (i) For which $a$ is $f$ convex? Give the smallest such $a$. (ii) For $a=0$, what is the global minimum value?

Only the symmetric part counts: $\tfrac12(A+A^\top)=\mathrm{diag}(2,a)$.

(i) Hessian $\mathrm{diag}(2,a)\succeq0\iff a\ge0$. (ii) With $a=0$: $f=x_1^2-2x_1$, independent of $x_2$. Minimum $-1$ at $x_1=1$, for every $x_2$: infinitely many global minimizers.

(Midterm 1, 2015 style.) $C$ is convex, $f$ is convex on $C$, and $\x_1\ne\x_2$ are both local minimizers in the interior of $C$. True or false: $\tfrac12(\x_1+\x_2)$ is a global minimizer.

For convex $f$, local = global, so $f(\x_1)=f(\x_2)=f^\star$. Now use the chord inequality.

True. Both are global with value $f^\star$, and $f(\tfrac12\x_1+\tfrac12\x_2)\le\tfrac12f(\x_1)+\tfrac12f(\x_2)=f^\star$, so the midpoint also achieves $f^\star$ (the set of minimizers is convex). No differentiability is needed.

  • Compare values at all critical points once you know a global minimum exists
  • Use coercivity (Weierstrass on a sublevel set) to justify existence
  • Symmetrize before testing convexity
  • Calling $f''=0$ points minima
  • Claiming a global minimum for a cubic
  • Testing convexity of $\x^\top A\x$ with a non-symmetric $A$'s eigenvalues
  1. Find critical points, classify with $f''$ or the Hessian, and resolve silent cases by sign analysis.
  2. Coercive ⇒ a global minimum exists among the critical points: compare values.
  3. Convex ⇒ local = global, and the minimizer set is convex.

True or false: if $f$ is coercive, then every local minimum is a global minimum.

True
Coercivity guarantees that a global minimum exists. Does it rule out other, higher dips?
False
$x^4-2x^2+x$ is coercive with two local minima at different heights.

True or false: if $f\in C^2$, $\grad f(\x^\star)=\0$ and $\hess f(\x^\star)\succ0$, then $\x^\star$ is a global minimum.

True
SOSC is a local statement. Think of a function with two dips.
False
SOSC gives a strict local minimum only; global needs more (e.g. convexity).

True or false: a convex function can have two distinct global minimizers.

True
$f(x_1,x_2)=x_1^2$ is minimized on the whole $x_2$-axis. Strict convexity is what forces uniqueness.
False
Convex is not strictly convex: think of a flat-bottomed valley.

Set B: descent, step sizes and iteration counts

The second half of the first midterm is about gradient methods: exact steps on quadratics, rates as fractions, and "how many iterations".

These questions reward knowing three formulas cold and spotting one trick per question.

Mental arithmetic in a shop: once you know the prices by heart, the bill adds itself up.

Lectures used: 6–10 (Parts 5–7).

(Midterm 1, 2015 style: watch the matrix.) $f(\x)=\x^\top\begin{pmatrix}1&-2\\2&2\end{pmatrix}\x+(2,4)\,\x+5$. Find $\x^\star$ and $f^\star$.

The off-diagonal entries cancel in $\x^\top M\x$: $f=x_1^2+2x_2^2+2x_1+4x_2+5$.

$f=x_1^2+2x_2^2+2x_1+4x_2+5$. Setting the gradient to zero: $2x_1+2=0$, $4x_2+4=0$, so $\x^\star=(-1,-1)$ and $f^\star=1+2-2-4+5=2$.

Same $f$. Steepest descent with exact line search from $\x_0=\0$. Using the guaranteed rate, find the smallest $N$ with $f(\x_N)-f^\star\le10^{-3}$.

Hessian $\mathrm{diag}(2,4)$, $\kappa=2$, $\rho=\big(\frac{\kappa-1}{\kappa+1}\big)^2$. And $f(\0)-f^\star=?$

$\rho=(1/3)^2=1/9$, $f(\0)-f^\star=5-2=3$. Need $3\cdot9^{-N}\le10^{-3}$, i.e. $9^N\ge3000$, $N\ge\ln3000/\ln9\approx3.64$, so $N=4$. (Running it gives gaps $0.222, 0.0165, 0.00122, 0.00009$: indeed step 4.)

(Midterm 1, 2024 style.) $f(\x)=\tfrac12\x^\top Q\x-\v^\top\x+4$, $Q\succ0$, $\kappa(Q)=10$, $Q\v=4\v$, $\norm{\v}=6$. Find $f^\star$ and the number of exact-line-search steepest-descent iterations from $\0$ to reach $\x^\star$.

$\x^\star=Q^{-1}\v=\v/4$. The condition number is a distractor: what direction is the first step?

$f^\star=4-\tfrac12\v^\top Q^{-1}\v=4-\tfrac{36}8=-\tfrac12$. $\grad f(\0)=-\v$, an eigenvector, so the line $\{\alpha\v\}$ passes through $\x^\star=\v/4$ and the exact step reaches it: 1 iteration.

(Midterm 1, 2017 style.) $f(\x)=\tfrac12\x^\top Q\x-\b^\top\x$, $Q=\begin{pmatrix}3&1\\1&3\end{pmatrix}$, $\b=(1,2)^\top$. Find $\x^\star$ and the smallest $L$ with $f\in C^1_L$.

Solve $Q\x=\b$. $L=\lambda_{\max}(Q)$; eigenvalues are $3\pm1$.

$Q^{-1}=\tfrac18\begin{pmatrix}3&-1\\-1&3\end{pmatrix}$, so $\x^\star=\tfrac18(1,5)=(0.125,0.625)$. $L=4$.

(Midterm 1, 2024 style.) $f\ge0$, $f\in C^1_L$ with $L=4$, $f(\x_0)=50$. (i) An inexact line search guarantees $f(\x_k)-f(\x_{k+1})\ge0.1\norm{\grad f(\x_k)}^2$. Find the smallest $T$ that guarantees some iterate with $\norm{\grad f}\le0.1$. (ii) With a constant step $\alpha$ and $\d=-\grad f$, the decrease is at least $\alpha(1-\tfrac{L\alpha}2)\norm{\grad f}^2$: give the best $\alpha$ and the new $T$.

Telescope: $c\sum_{k\lt T}\norm{\g_k}^2\le f(\x_0)-f^\star\le f(\x_0)$, so $\min\norm{\g_k}^2\le f(\x_0)/(cT)$.

(i) Need $50/(0.1T)\le0.01$: $T\ge50000$. (ii) $\alpha(1-2\alpha)$ is maximized at $\alpha=1/4=1/L$, giving $c=\tfrac18$; then $50/(T/8)\le0.01$: $T\ge40000$.

(Midterm 1, 2017 style.) Minimizing a $C^1$ function with steepest descent, you observe $\grad f(\x_{k+1})^\top\grad f(\x_k)=0$ for every $k$. Which step rule was used?

What is $\phi'(\alpha_k)$ for $\phi(\alpha)=f(\x_k-\alpha\g_k)$?

$\phi'(\alpha_k)=-\grad f(\x_{k+1})^\top\g_k=0$: the step is a stationary point of $\phi$, which is what exact line search produces.

$f$ is $\mu$-strongly convex and $L$-smooth with $\mu=2$, $L=18$. Gradient descent with $h=2/(\mu+L)$. Give $h$, the contraction factor of $\norm{\x_k-\x^\star}$, and the fewest iterations guaranteeing $\norm{\x_k-\x^\star}\le10^{-3}\norm{\x_0-\x^\star}$.

$Q_f=9$, factor $(Q_f-1)/(Q_f+1)$.

$h=0.1$, factor $0.8$. $0.8^k\le10^{-3}$ needs $k\ge\ln1000/\ln1.25\approx30.96$: $k=31$.

$f(x)=x^2$ at $x=1$ with $d=-f'(1)=-2$ and $c_1=\tfrac12$. What is the largest step $\alpha$ satisfying Armijo?

Armijo: $(1-2\alpha)^2\le1+c_1\alpha\cdot f'(1)d=1-2\alpha$.

$1-4\alpha+4\alpha^2\le1-2\alpha\iff4\alpha^2\le2\alpha\iff\alpha\le\tfrac12$. (And $\alpha=\tfrac12$ is the exact minimizer here: $c_1=\tfrac12$ is the largest $c_1$ that still accepts it.)

  • Check whether the first gradient is an eigenvector before computing anything
  • Use $f(\x_0)-f^\star\le f(\x_0)$ when only $f\ge0$ is given
  • Quote $\rho$ for $f$-gaps as $\big(\frac{\kappa-1}{\kappa+1}\big)^2$
  • Using the non-symmetric matrix to get $\x^\star$
  • Confusing the distance factor $\frac{\kappa-1}{\kappa+1}$ with the $f$-gap factor (its square)
  • Rounding $N$ down
  1. Exact line search on quadratics: $\rho=\big(\frac{\kappa-1}{\kappa+1}\big)^2$ for $f$-gaps; one step if the start gradient is an eigenvector.
  2. Iteration counts come from telescoping a per-step decrease $c\norm{\g}^2$.
  3. Orthogonal successive gradients is the fingerprint of exact line search.

True or false: steepest descent with exact line search on a convex quadratic whose Hessian eigenvalues are all equal converges in one iteration from any start.

True
$Q=\lambda I$: $-\grad f$ points straight at $\x^\star$, and the exact step lands on it.
False
With all eigenvalues equal, the level sets are circles. Where does $-\grad f$ point?

True or false: for any descent direction with any line search, there is $C\gt0$ with $f(\x_k)-f(\x_{k+1})\ge C\norm{\grad f(\x_k)}^2$.

True
Consider directions almost perpendicular to the gradient, or tiny steps.
False
It needs both a step rule (e.g. Armijo–Goldstein or Wolfe) and an angle condition $\cos\theta_k\ge\delta$.

If $\kappa=5$, exact-line-search descent on a quadratic has $f$-gap factor…

$\tfrac23$
That's $\frac{\kappa-1}{\kappa+1}$, the distance factor. Function gaps get the square.
$\tfrac49$
$(4/6)^2=4/9$, exactly the 2024 midterm's boxed answer.
$\tfrac15$
The rate depends on $\frac{\kappa-1}{\kappa+1}$, not on $1/\kappa$.

Set C: conjugate gradients, Newton and quasi-Newton

Questions on Lectures 11–14 test structure: how many CG steps, what a secant update must satisfy, and when positive definiteness can be lost.

They look intimidating but usually reduce to one theorem and a two-line computation.

A combination lock: once you know which dial controls what, it opens quickly.

Lectures used: 11–14 (Parts 8–10).

(Midterm 2, 2017 style.) $Q$ is $10\times10$, symmetric positive definite, with distinct eigenvalues and unit eigenvectors $\e_i$. $\b=2\e_1+3\e_4$ and $\x_0=\0$. How many CG iterations solve $Q\x=\b$?

The initial gradient $-\b$ lives in the span of two eigenvectors, and so does every Krylov vector.

$\mathcal K$ stays inside $\mathrm{span}\{\e_1,\e_4\}$, and on that 2-D subspace $Q$ has 2 distinct eigenvalues: a degree-2 polynomial vanishing at $\lambda_1,\lambda_4$ kills the error. 2 iterations.

(Midterm 2, 2015 style.) $A=I+3\sum_{i=1}^4\u_i\u_i^\top$ on $\R^{50}$, with $\u_i$ orthonormal. At most how many CG iterations solve $A\x=\b$ from any start?

What are the eigenvalues of $A$? (Look at $A\u_i$ and at vectors orthogonal to all $\u_i$.)

$A\u_i=4\u_i$ and $A\v=\v$ for $\v\perp$ all $\u_i$: only two distinct eigenvalues, 4 and 1. So at most 2 iterations.

$A=\begin{pmatrix}2&1\\1&3\end{pmatrix}$, $\d_0=(1,0)^\top$. For which $t$ is $\d_1=(t,1)^\top$ conjugate to $\d_0$?

Need $\d_0^\top A\d_1=0$; $\d_0^\top A=(2,1)$.

$2t+1=0$, so $t=-\tfrac12$.

CG on $f=\tfrac12\x^\top A\x-\b^\top\x$ with $A=\begin{pmatrix}4&1\\1&3\end{pmatrix}$, $\b=(1,2)^\top$, $\x_0=\0$. Find $\alpha_0$ and $\beta_1$.

$\g_0=-\b$, $\alpha_0=\norm{\g_0}^2/\g_0^\top A\g_0$, $\x_1=\x_0+\alpha_0\b$, $\beta_1=\norm{\g_1}^2/\norm{\g_0}^2$.

$\norm{\g_0}^2=5$, $\g_0^\top A\g_0=20$, so $\alpha_0=\tfrac14$ and $\x_1=(\tfrac14,\tfrac12)$. $\g_1=A\x_1-\b=(\tfrac12,-\tfrac14)$, $\norm{\g_1}^2=\tfrac5{16}$, so $\beta_1=\tfrac1{16}$.

Newton's method for minimizing $f(x)=x^4$ from $x_0\ne0$. Each step multiplies $x_k$ by a constant. Which?

$x_{k+1}=x_k-f'(x_k)/f''(x_k)=x_k-4x_k^3/(12x_k^2)$.

$x_{k+1}=\tfrac23x_k$: only linear convergence, because $f''(0)=0$ breaks the quadratic-convergence hypothesis.

Newton's root-finding iteration for $\phi(t)=t/\sqrt{1+t^2}$ is $t_{k+1}=-t_k^3$ ([Y] Example 1.2.4). From $t_0=\tfrac12$, find $t_2$.

$t_1=-\tfrac18$.

$t_1=-(1/2)^3=-1/8$, $t_2=-(-1/8)^3=1/512\approx0.00195$. It converges because $|t_0|\lt1$; from $|t_0|\gt1$ it diverges.

SR1 update from $H_0=I$ with $\boldsymbol\delta=(1,2)^\top$, $\boldsymbol\gamma=(2,5)^\top$. Compute $H_1$ (it must satisfy $H_1\boldsymbol\gamma=\boldsymbol\delta$).

$\u=\boldsymbol\delta-H_0\boldsymbol\gamma=(-1,-3)$ and $H_1=H_0+\u\u^\top/(\u^\top\boldsymbol\gamma)$.

$\u^\top\boldsymbol\gamma=-17$, so $H_1=I-\tfrac1{17}\begin{pmatrix}1&3\\3&9\end{pmatrix}=\tfrac1{17}\begin{pmatrix}16&-3\\-3&8\end{pmatrix}$. Check: $H_1\boldsymbol\gamma=\tfrac1{17}(17,34)=(1,2)$.

(Midterm 1, 2015 style.) Testing a rank-1 (SR1) quasi-Newton code, you see $G^{(1)},G^{(3)}$ positive definite but $G^{(2)},G^{(4)}$ with negative eigenvalues. Can you conclude there's a bug?

What sign can the SR1 denominator $(\boldsymbol\delta-H\boldsymbol\gamma)^\top\boldsymbol\gamma$ take?

No. The SR1 denominator can be negative, so a rank-1 correction can subtract and make the matrix indefinite; SR1 does not preserve positive definiteness (unlike DFP/BFGS with $\boldsymbol\delta^\top\boldsymbol\gamma\gt0$).

  • Count distinct eigenvalues "seen" by $\b$ (or the start error) for CG
  • Verify a quasi-Newton update by checking the secant condition
  • State which property (PD, secant, hereditary) each update guarantees
  • Answering "$n$ iterations" for CG when the spectrum or $\b$ is special
  • Assuming every quasi-Newton update keeps $H\succ0$
  • Expecting quadratic convergence of Newton at a degenerate minimizer
  1. CG needs as many steps as distinct eigenvalues present in the starting error.
  2. Secant condition $H_{k+1}\boldsymbol\gamma_k=\boldsymbol\delta_k$ is the check for any quasi-Newton update.
  3. Newton is quadratic only near a non-degenerate minimizer; SR1 may lose positive definiteness.

True or false: in exact arithmetic, CG solves any $n\times n$ SPD system in at most $n$ iterations.

True
Conjugate directions are independent, and the expanding subspace reaches all of $\R^n$.
False
Recall the conjugate direction theorem.

True or false: Newton's direction $-[\hess f]^{-1}\grad f$ is always a descent direction.

True
What if the Hessian is indefinite, say near a saddle?
False
It's a descent direction when $\hess f\succ0$; otherwise it can point uphill (or towards a saddle).

DFP with exact line search on an $n$-dimensional SPD quadratic terminates in at most…

$n$ iterations, with $H_n=G^{-1}$
Quadratic termination: the directions are conjugate and the hereditary property builds $G^{-1}$.
1 iteration
That's Newton (or a quasi-Newton method that starts with $H_0=G^{-1}$).
$\sqrt\kappa$ iterations
$\sqrt\kappa$ appears in CG's rate bound, not in finite termination.

Every continuous function on a closed set attains its minimum.

True
Weierstrass needs bounded too. Try $e^x$ on $(-\infty,0]$.
False
$e^x$ on the closed set $(-\infty,0]$ has infimum 0, never attained.

If $\grad f(\x^\star)=\0$ and $\hess f(\x^\star)\succeq0$, then $\x^\star$ is a local minimum.

True
Semidefinite leaves the test silent. Try $x^3$ at 0.
False
$f=x^3$: $f'(0)=0$, $f''(0)=0\succeq0$, but 0 is not a minimum.

For a convex $C^1$ function, $\grad f(\x)=\0$ implies $\x$ is a global minimizer.

True
Tangent-plane inequality: $f(\y)\ge f(\x)+\0^\top(\y-\x)$.
False
Write the first-order convexity inequality at $\x$.

The Armijo condition alone prevents steps that are too short.

True
Armijo accepts every small enough step. What rules those out?
False
Tiny steps satisfy Armijo; the curvature (Wolfe) condition or backtracking from a fixed $\bar\alpha$ handles that.

For convex $L$-smooth $f$, gradient descent with $h=1/L$ gives $f(\x_k)-f^\star=O(1/k)$.

True
[Y] Thm 2.1.14: $\le2L\norm{\x_0-\x^\star}^2/(k+4)$.
False
That's exactly the content of Lecture 9's theorem.

If an SPD matrix has 3 distinct eigenvalues, CG converges in at most 3 iterations.

True
A degree-3 polynomial with $Q(0)=1$ can vanish on all three eigenvalues.
False
Think about the min–max polynomial bound.

BFGS keeps $H_{k+1}\succ0$ whenever $H_k\succ0$, regardless of the step.

True
There's one condition on $\boldsymbol\delta_k$ and $\boldsymbol\gamma_k$.
False
It needs $\boldsymbol\delta_k^\top\boldsymbol\gamma_k\gt0$, which Wolfe or exact line searches guarantee.
Every question, answered right

You made it all the way down.

You worked through every quiz and every practice problem in Downhill. That took real effort, and the understanding is yours to keep. Add your name to make a certificate.