diff --git a/lectures/BCG_complete_mkts.md b/lectures/BCG_complete_mkts.md index 03163c21..14308236 100644 --- a/lectures/BCG_complete_mkts.md +++ b/lectures/BCG_complete_mkts.md @@ -27,17 +27,14 @@ In addition to what's in Anaconda, this lecture will need the following librarie tags: [hide-output] --- !pip install --upgrade quantecon -!conda install -y -c plotly plotly plotly-orca +!pip install kaleido ``` ## Introduction -This is a prolegomenon to another lecture {doc}`Equilibrium Capital Structures with Incomplete Markets ` about a model with -incomplete markets authored by Bisin, Clementi, and Gottardi {cite}`BCG_2018`. +This is a prolegomenon to another lecture {doc}`Equilibrium Capital Structures with Incomplete Markets ` about a model with incomplete markets authored by Bisin, Clementi, and Gottardi {cite}`BCG_2018`. -We adopt specifications of preferences and technologies very close to -Bisin, Clemente, and Gottardi’s but unlike them assume that there are complete -markets in one-period Arrow securities. +We adopt specifications of preferences and technologies very close to Bisin, Clementi, and Gottardi’s but, unlike them, assume that there are complete markets in one-period Arrow securities. This simplification of BCG’s setup helps us by @@ -53,8 +50,7 @@ This simplification of BCG’s setup helps us by - introducing `Big K, little k` issues in a simple context that will recur in the BCG incomplete markets environment -A Big K, little k analysis also played roles in [this quantecon lecture](https://python.quantecon.org/cass_koopmans_1.html) as well as -[here](https://python.quantecon.org/rational_expectations.html) and {doc}`here `. +A Big K, little k analysis also played roles in [this quantecon lecture](https://python.quantecon.org/cass_koopmans_1.html) as well as [here](https://python.quantecon.org/rational_expectations.html) and {doc}`here `. ### Setup @@ -64,18 +60,17 @@ There are two types of consumers named $i=1,2$. A scalar random variable $\epsilon$ with probability density $g(\epsilon)$ affects both -- the return in period $1$ from investing - $k \geq 0$ in physical capital in period $0$. +- the return in period $1$ from investing + $k \geq 0$ in physical capital in period $0$ - exogenous period $1$ endowments of the consumption good for agents of types $i =1$ and $i=2$. -Type $i=1$ and $i=2$ agents’ period $1$ endowments are -correlated with the return on physical capital in different ways. +Type $i=1$ and $i=2$ agents’ period $1$ endowments are correlated with the return on physical capital in different ways. We discuss two arrangements: - a command economy in which a benevolent planner chooses $k$ and - allocates goods to the two types of consumers in each period and each random + allocates goods to the two types of consumers in each period and each random second period state - a competitive equilibrium with markets in claims on physical capital and a complete set (possibly a continuum) of one-period Arrow @@ -84,8 +79,7 @@ We discuss two arrangements: ### Endowments -There is a single consumption good in period $0$ and at each -random state $\epsilon$ in period $1$. +There is a single consumption good in period $0$ and at each random state $\epsilon$ in period $1$. Economy-wide endowments in periods $0$ and $1$ are @@ -96,12 +90,11 @@ w_1(\epsilon) & \textrm{ in state }\epsilon \end{aligned} $$ -Soon we’ll explain how aggregate endowments are divided between -type $i=1$ and type $i=2$ consumers. +Soon we’ll explain how aggregate endowments are divided between type $i=1$ and type $i=2$ consumers. We don’t need to do that in order to describe a social planning problem. -### Technology: +### Technology Where $\alpha \in (0,1)$ and $A >0$ @@ -112,11 +105,9 @@ $$ \end{aligned} $$ -### Preferences: +### Preferences -A consumer of type $i$ orders period $0$ consumption -$c_0^i$ and state $\epsilon$, period $1$ consumption -$c^i_1(\epsilon)$ by +A consumer of type $i$ orders period $0$ consumption $c_0^i$ and state $\epsilon$, period $1$ consumption $c^i_1(\epsilon)$ by $$ u^i = u(c_0^i) + \beta \int u(c_1^i(\epsilon)) g (\epsilon) d \epsilon, \quad i = 1,2 @@ -139,14 +130,11 @@ $$ \begin{aligned} \epsilon & \sim {\mathcal N}(\mu, \sigma^2) \cr u(c) & = \frac{c^{1-\gamma}}{1 - \gamma} \cr -w_1^i(\epsilon) & = e^{- \chi_i \mu - .5 \chi_i^2 \sigma^2 + \chi_i \epsilon} , \quad \chi_i \in [0,1] +w_1^i(\epsilon) & = e^{- \chi_i \mu - .5 \chi_i^2 \sigma^2 + \chi_i \epsilon} , \quad \chi_i \in [-1,1] \end{aligned} $$ -Sometimes instead of asuming $\epsilon \sim g(\epsilon) = {\mathcal N}(0,\sigma^2)$, -we’ll assume that $g(\cdot)$ is a probability -mass function that serves as a discrete approximation to a standardized -normal density. +Sometimes instead of assuming $\epsilon \sim g(\epsilon) = {\mathcal N}(\mu,\sigma^2)$, we’ll assume that $g(\cdot)$ is a probability mass function that serves as a discrete approximation to this normal density. ### Pareto criterion and planning problem @@ -156,8 +144,7 @@ $$ \textrm{obj} = \phi_1 u^1 + \phi_2 u^2 , \quad \phi_i \geq 0, \quad \phi_1 + \phi_2 = 1 $$ -where $\phi_i \geq 0$ is a Pareto weight that the planner attaches -to a consumer of type $i$. +where $\phi_i$ is the Pareto weight that the planner attaches to a consumer of type $i$. We form the following Lagrangian for the planner’s problem: @@ -185,14 +172,12 @@ The first four equations imply that $$ \begin{aligned} -\frac{u'(c_1^1(\epsilon))}{u'(c_0^1))} & = \frac{u'(c_1^2(\epsilon))}{u'(c_0^2))} = \frac{\lambda_1(\epsilon)}{\lambda_0} \cr +\frac{u'(c_1^1(\epsilon))}{u'(c_0^1)} & = \frac{u'(c_1^2(\epsilon))}{u'(c_0^2)} = \frac{\lambda_1(\epsilon)}{\lambda_0} \cr \frac{u'(c_0^1)}{u'(c_0^2)} & = \frac{u'(c_1^1(\epsilon))}{u'(c_1^2(\epsilon))} = \frac{\phi_2}{\phi_1} \end{aligned} $$ -These together with the fifth first-order condition for the planner -imply the following equation that determines an optimal choice of -capital +These together with the fifth first-order condition for the planner imply the following equation that determines an optimal choice of capital $$ 1 = \beta \alpha A k^{\alpha -1} \int \frac{u'(c_1^i(\epsilon))}{u'(c_0^i)} e^\epsilon g(\epsilon) d \epsilon @@ -214,8 +199,7 @@ $$ \frac{u'(c^1)}{u'(c^2)} = \left(\frac{c^1}{c^2}\right)^{-\gamma} = \frac{\phi_2}{\phi_1} $$ -where it is to be understood that this equation holds for $c^1 = c^1_0$ and $c^2 = c^2_0$ and also -for $c^1 = c^1(\epsilon)$ and $c^2 = c^2(\epsilon)$ for all $\epsilon$. +where it is to be understood that this equation holds for $c^1 = c^1_0$ and $c^2 = c^2_0$ and also for $c^1 = c^1_1(\epsilon)$ and $c^2 = c^2_1(\epsilon)$ for all $\epsilon$. With the same understanding, it follows that @@ -234,11 +218,9 @@ $$ \end{aligned} $$ -where $\eta \in [0,1]$ is a function of $\phi_1$ and -$\gamma$. +where $\eta \in [0,1]$ is a function of $\phi_1$ and $\gamma$. -Consequently, we can write the planner’s first-order condition for -$k$ as +Consequently, we can write the planner’s first-order condition for $k$ as $$ 1 = \beta \alpha A k^{\alpha -1} \int \left( \frac{w_1(\epsilon) + A k^\alpha e^\epsilon} @@ -247,9 +229,7 @@ $$ which is one equation to be solved for $k \geq 0$. -Anticipating a `Big K, little k` idea widely used in macroeconomics, -to be discussed in detail below, let $K$ be the value of $k$ -that solves the preceding equation so that +Anticipating a `Big K, little k` idea widely used in macroeconomics, to be discussed in detail below, let $K$ be the value of $k$ that solves the preceding equation so that ```{math} :label: focke @@ -271,39 +251,29 @@ c_1^2 (\epsilon) & = (1 - \eta) C_1(\epsilon) \end{aligned} $$ -where $\eta \in [0,1]$ is the consumption share parameter -mentioned above that is a function of the Pareto weight $\phi_1$ -and the utility curvature parameter $\gamma$. +where $\eta \in [0,1]$ is the consumption share parameter mentioned above that is a function of the Pareto weight $\phi_1$ and the utility curvature parameter $\gamma$. #### Remarks -The relative Pareto weight parameter $\eta$ does not appear in -equation {eq}`focke` that determines $K$. +The consumption share parameter $\eta$, which is determined by the Pareto weights, does not appear in equation {eq}`focke` that determines $K$. -Neither does it influence $C_0$ or $C_1(\epsilon)$, which -depend solely on $K$. +Neither does it influence $C_0$ or $C_1(\epsilon)$, which depend solely on $K$. -The role of $\eta$ is to determine how to allocate total -consumption between the two types of consumers. +The role of $\eta$ is to determine how to allocate total consumption between the two types of consumers. -Thus, the planner’s choice of $K$ does not interact with how it wants to allocate consumption. +Thus, the planner’s choice of $K$ does not interact with how it wants to allocate consumption. ## Competitive equilibrium -We now describe a competitive equilibrium for an economy that has -specifications of consumer preferences, technology, and aggregate -endowments that are identical to those in the preceding planning -problem. +We now describe a competitive equilibrium for an economy that has specifications of consumer preferences, technology, and aggregate endowments that are identical to those in the preceding planning problem. -While prices do not appear in the planning problem – only quantities do – -prices play an important role in a competitive equilibrium. +While prices do not appear in the planning problem – only quantities do – prices play an important role in a competitive equilibrium. -To understand how the planning economy is related to a competitive -equilibrium, we now turn to the `Big K, little k` distinction. +To understand how the planning economy is related to a competitive equilibrium, we now turn to the `Big K, little k` distinction. ### Measures of agents and firms -We follow BCG in assuming that there are unit measures of +We follow BCG in assuming that there are unit measures of - consumers of type $i=1$ - consumers of type $i=2$ @@ -312,8 +282,7 @@ We follow BCG in assuming that there are unit measures of $A k^\alpha e^\epsilon$ units of the time $1$ good in random state $\epsilon$ -Thus, let $\omega \in [0,1]$ index a particular consumer of type -$i$. +Thus, let $\omega \in [0,1]$ index a particular consumer of type $i$. Then define Big $C^i$ as @@ -322,33 +291,29 @@ C^i = \int_0^1 c^i(\omega) d \, \omega $$ In the same spirit, let $\zeta \in [0,1]$ index a particular firm. + Then define Big $K$ as $$ K = \int_0^1 k(\zeta) d \, \zeta $$ -The assumption that there are continua of our three types of -agents plays an important role making each individual agent into a -powerless **price taker**: +The assumption that there are continua of our three types of agents plays an important role in making each individual agent into a powerless **price taker**: -- an individual consumer chooses its own (infinesimal) part +- an individual consumer chooses its own (infinitesimal) part $c^i(\omega)$ of $C^i$ taking prices as given -- an individual firm chooses its own (infinitesmimal) part - $k(\zeta)$ of $K$ taking prices as +- an individual firm chooses its own (infinitesimal) part + $k(\zeta)$ of $K$ taking prices as given - equilibrium prices depend on the `Big K, Big C` objects $K$ and $C$ -Nevertheless, in equilibrium, $K = k, C^i = c^i$ +Nevertheless, in equilibrium, $K = k$ and $C^i = c^i$. -The assumption about measures of agents is thus a powerful device for -making a host of competitive agents take as given equilibrium prices -that are determined by the independent decisions of hosts of agents who behave just like they do. +The assumption about measures of agents is thus a powerful device for making a host of competitive agents take as given equilibrium prices that are determined by the independent decisions of hosts of agents who behave just like they do. #### Ownership -Consumers of type $i$ own the following exogenous quantities of -the consumption good in periods $0$ and $1$: +Consumers of type $i$ own the following exogenous quantities of the consumption good in periods $0$ and $1$: $$ \begin{aligned} @@ -366,14 +331,9 @@ $$ \end{aligned} $$ -Consumers also own shares in a firm that operates the technology for converting -nonnegative amounts of the time $0$ consumption good one-for-one -into a capital good $k$ that produces -$A k^\alpha e^\epsilon$ units of the time $1$ consumption good -in time $1$ state $\epsilon$. +Consumers also own shares in a firm that operates the technology for converting nonnegative amounts of the time $0$ consumption good one-for-one into a capital good $k$ that produces $A k^\alpha e^\epsilon$ units of the time $1$ consumption good in time $1$ state $\epsilon$. -Consumers of types $i=1,2$ are endowed with $\theta_0^i$ -shares of a firm and +Consumers of types $i=1,2$ are endowed with $\theta_0^i$ shares of a firm and $$ \theta_0^1 + \theta_0^2 = 1 @@ -381,25 +341,23 @@ $$ #### Asset markets -At time $0$, consumers trade the following assets with other consumers -and with firms: +At time $0$, consumers trade the following assets with other consumers and with firms: - equities (also known as stocks) issued by firms - one-period Arrow securities that pay one unit of consumption at time $1$ when the shock $\epsilon$ assumes a particular value -Later, we’ll allow the firm to issue bonds too, but -not now. +Later, we’ll allow the firm to issue bonds too, but not now. ### Objects appearing in a competitive equilibrium Let -- $a^i(\epsilon)$ be consumer $i$ ’s purchases of claims +- $a^i(\epsilon)$ be consumer $i$’s purchases of claims on time $1$ consumption in state $\epsilon$ - $q(\epsilon)$ be a pricing kernel for one-period Arrow securities -- $\theta_0^i \geq 0$ be consumer $i$'s intial share of +- $\theta_0^i \geq 0$ be consumer $i$'s initial share of the firm, $\sum_i \theta_0^i =1$ - $\theta^i$ be the fraction of a firm’s shares purchased by consumer $i$ at time $t=0$ @@ -407,33 +365,26 @@ Let - $\tilde V$ be the value of equity issued by the representative firm - $K, C_0$ be two scalars and $C_1(\epsilon)$ a function - that we use to construct a guess about an equilibrium pricing kernel + that we use to construct a guess about an equilibrium pricing kernel for Arrow securities -We proceed to describe constrained optimum problems faced by -consumers and a representative firm in a competitive equilibrium. +We proceed to describe constrained optimum problems faced by consumers and a representative firm in a competitive equilibrium. ### A representative firm’s problem -A representative firm takes Arrow security prices $q(\epsilon)$ as -given. +A representative firm takes Arrow security prices $q(\epsilon)$ as given. -The firm purchases capital $k \geq 0$ from consumers at time -$0$ and finances itself by issuing equity at time $0$. +The firm purchases capital $k \geq 0$ from consumers at time $0$ and finances itself by issuing equity at time $0$. -The firm produces time $1$ goods $A k^\alpha e^\epsilon$ in -state $\epsilon$ and pays all of these `earnings` to owners of its -equity. +The firm produces time $1$ goods $A k^\alpha e^\epsilon$ in state $\epsilon$ and pays all of these `earnings` to owners of its equity. -The value of a firm's equity at time $0$ can be computed by multiplying -its state-contingent earnings by their Arrow securities prices and then -adding over all contingencies: +The value of a firm's equity at time $0$ can be computed by multiplying its state-contingent earnings by their Arrow securities prices and then adding over all contingencies: $$ \tilde V = \int A k^\alpha e^\epsilon q(\epsilon) d \epsilon $$ -Owners of a firm want it to choose $k$ to maximize +Owners of a firm want it to choose $k$ to maximize $$ V = - k + \int A k^\alpha e^\epsilon q(\epsilon) d \epsilon @@ -451,19 +402,15 @@ $$ V = - k + \tilde V $$ -The right side equals the value of equity minus the cost of the time $0$ goods -that it purchases and uses as capital. +The right side equals the value of equity minus the cost of the time $0$ goods that it purchases and uses as capital. ### A consumer’s problem We now pose a consumer’s problem in a competitive equilibrium. -As a price taker, each consumer faces a given Arrow securities pricing kernel -$q(\epsilon)$, a given value of a firm $V$ that has chosen capital stock $k$, a price of -equity $\tilde V$, and prospective next period random dividends $A k^\alpha e^\epsilon$. +As a price taker, each consumer faces a given Arrow securities pricing kernel $q(\epsilon)$, a given value of a firm $V$ that has chosen capital stock $k$, a price of equity $\tilde V$, and prospective next period random dividends $A k^\alpha e^\epsilon$. -If we evaluate consumer $i$'s time $1$ budget constraint at zero consumption $c^i_1(\epsilon) = 0$ and solve for $-a^i(\epsilon)$ -we obtain +If we evaluate consumer $i$'s time $1$ budget constraint at zero consumption $c^i_1(\epsilon) = 0$ and solve for $-a^i(\epsilon)$ we obtain ```{math} :label: debtlimit @@ -471,24 +418,22 @@ we obtain -\bar a^i(\epsilon;\theta^i) = w_1^i(\epsilon) +\theta^i A k^\alpha e^\epsilon ``` -The quantity $- \bar a^i(\epsilon;\theta^i)$ is the maximum amount that it is feasible for consumer $i$ to repay to -his Arrow security creditors at time $1$ in state $\epsilon$. +The quantity $- \bar a^i(\epsilon;\theta^i)$ is the maximum amount that it is feasible for consumer $i$ to repay to his Arrow security creditors at time $1$ in state $\epsilon$. Notice that $-\bar a^i(\epsilon;\theta^i)$ defined in {eq}`debtlimit` depends on * his endowment $w_1^i(\epsilon)$ at time $1$ in state $\epsilon$ -* his share $\theta^i$ of a representive firm's dividends +* his share $\theta^i$ of a representative firm's dividends -These constitute two sources of **collateral** that back the consumer's issues of Arrow securities that pay off in state $\epsilon$ +These constitute two sources of **collateral** that back the consumer's issues of Arrow securities that pay off in state $\epsilon$. -Consumer $i$ chooses a scalar $c_0^i$ and a function -$c_1^i(\epsilon)$ to maximize +Consumer $i$ chooses a scalar $c_0^i$ and a function $c_1^i(\epsilon)$ to maximize $$ u(c_0^i) + \beta \int u(c_1^i(\epsilon)) g (\epsilon) d \epsilon $$ -subject to time $0$ and time $1$ budget constraints +subject to time $0$ and time $1$ budget constraints $$ \begin{aligned} @@ -497,35 +442,29 @@ c_1^i(\epsilon) & \leq w_1^i(\epsilon) +\theta^i A k^\alpha e^\epsilon + a^i(\ep \end{aligned} $$ -Attach Lagrange multiplier $\lambda_0^i$ to the budget constraint -at time $0$ and scaled Lagrange multiplier -$\beta \lambda_1^i(\epsilon) g(\epsilon)$ to the budget constraint -at time $1$ and state $\epsilon$, -then form the Lagrangian +Attach Lagrange multiplier $\lambda_0^i$ to the budget constraint at time $0$ and scaled Lagrange multiplier $\beta \lambda_1^i(\epsilon) g(\epsilon)$ to the budget constraint at time $1$ and state $\epsilon$, then form the Lagrangian $$ \begin{aligned} L^i & = u(c_0^i) + \beta \int u(c^i_1(\epsilon)) g(\epsilon) d \epsilon \cr - & + \lambda_0^i [ w_0^i + \theta_0^i - \int q(\epsilon) a^i(\epsilon) d \epsilon - + & + \lambda_0^i [ w_0^i + \theta_0^i V - \int q(\epsilon) a^i(\epsilon) d \epsilon - \theta^i \tilde V - c_0^i ] \cr & + \beta \int \lambda_1^i(\epsilon) [ w_1^i(\epsilon) + \theta^i A k^\alpha e^\epsilon - + a^i(\epsilon) c_1^i(\epsilon) ] g(\epsilon) d \epsilon + + a^i(\epsilon) - c_1^i(\epsilon) ] g(\epsilon) d \epsilon \end{aligned} $$ -Off corners, first-order necessary conditions for an optimum with respect to -$c_0^i, c_1^i(\epsilon),$ and $a^i(\epsilon)$ are +Off corners, first-order necessary conditions for an optimum with respect to $c_0^i, c_1^i(\epsilon),$ and $a^i(\epsilon)$ are $$ \begin{aligned} c_0^i: \quad & u'(c_0^i) - \lambda_0^i = 0 \cr c_1^i(\epsilon): \quad & \beta u'(c_1^i(\epsilon)) g(\epsilon) - \beta \lambda_1^i(\epsilon) g(\epsilon) = 0 \cr -a^i(\epsilon): \quad & -\lambda_0^i q(\epsilon) + \beta \lambda_1^i(\epsilon) = 0 +a^i(\epsilon): \quad & -\lambda_0^i q(\epsilon) + \beta \lambda_1^i(\epsilon) g(\epsilon) = 0 \end{aligned} $$ -These equations imply that consumer $i$ adjusts its consumption -plan to satisfy +These equations imply that consumer $i$ adjusts its consumption plan to satisfy ```{math} :label: qgeqn @@ -533,17 +472,13 @@ plan to satisfy q(\epsilon) = \beta \left( \frac{u'(c_1^i(\epsilon))}{u'(c_0^i)} \right) g(\epsilon) ``` -To deduce a restriction on equilibrium prices, we -solve the period $1$ budget constraint to express -$a^i(\epsilon)$ as +To deduce a restriction on equilibrium prices, we solve the period $1$ budget constraint to express $a^i(\epsilon)$ as $$ a^i(\epsilon) = c_1^i(\epsilon) - w_1^i(\epsilon) - \theta^i A k^\alpha e^\epsilon $$ -then substitute the expression on the right side into the time $0$ -budget constraint and rearrange to get the single intertemporal budget -constraint +then substitute the expression on the right side into the time $0$ budget constraint and rearrange to get the single intertemporal budget constraint ```{math} :label: noarb @@ -552,23 +487,17 @@ w_0^i + \theta_0^i V + \int w_1^i(\epsilon) q(\epsilon) d \epsilon + \theta^i \l \geq c_0^i + \int c_1^i(\epsilon) q(\epsilon) d \epsilon ``` -The right side of inequality {eq}`noarb` is the present value -of consumer $i$’s consumption while the left side is the present -value of consumer $i$’s endowment when consumer $i$ buys -$\theta^i$ shares of equity. +The right side of inequality {eq}`noarb` is the present value of consumer $i$’s consumption while the left side is the present value of consumer $i$’s endowment when consumer $i$ buys $\theta^i$ shares of equity. -From inequality {eq}`noarb`, we deduce two -findings. +From inequality {eq}`noarb`, we deduce two findings. -**1. No arbitrage profits condition:** +**1. No-arbitrage condition** Unless -```{math} -:label: tilde - +$$ \tilde V = A k^\alpha \int e^\epsilon q (\epsilon) d \epsilon -``` +$$ an **arbitrage** opportunity would be open. @@ -578,8 +507,7 @@ $$ \tilde V > A k^\alpha \int e^\epsilon q (\epsilon) d \epsilon $$ -the consumer could afford an arbitrarily high present value of consumption by setting $\theta^i$ to an arbitrarily large **negative** -number. +the consumer could afford an arbitrarily high present value of consumption by setting $\theta^i$ to an arbitrarily large **negative** number. If @@ -587,15 +515,11 @@ $$ \tilde V < A k^\alpha \int e^\epsilon q (\epsilon) d \epsilon $$ -the consumer could afford an arbitrarily high present value of -consumption by setting $\theta^i$ to be arbitrarily large **positive** -number. +the consumer could afford an arbitrarily high present value of consumption by setting $\theta^i$ to an arbitrarily large **positive** number. -Since resources are finite, there can exist no such arbitrage -opportunity in a competitive equilibrium. +Since resources are finite, there can exist no such arbitrage opportunity in a competitive equilibrium. -Therefore, it must be true -that the following no arbitrage condition prevails: +Therefore, it must be true that the following no-arbitrage condition prevails: ```{math} :label: tildeV20 @@ -603,50 +527,39 @@ that the following no arbitrage condition prevails: \tilde V = \int A k^\alpha e^\epsilon q(\epsilon;K) d \epsilon ``` -Equation {eq}`tildeV20` asserts that the value of equity -equals the value of the state-contingent dividends -$Ak^\alpha e^\epsilon$ evaluated at the Arrow security prices -$q(\epsilon; K)$ that we have expressed as a function of $K$. +Equation {eq}`tildeV20` asserts that the value of equity equals the value of the state-contingent dividends $Ak^\alpha e^\epsilon$ evaluated at the Arrow security prices $q(\epsilon; K)$, which we shall express as a function of $K$ in {eq}`arrowprices` below. We'll say more about this equation later. **2. Indeterminacy of portfolio** -When the no-arbitrage pricing equation {eq}`tildeV20` -prevails, a consumer of type $i$’s choice $\theta^i$ of equity is -indeterminate. +When the no-arbitrage pricing equation {eq}`tildeV20` prevails, a consumer of type $i$’s choice $\theta^i$ of equity is indeterminate. -Consumer of type $i$ can offset any choice of -$\theta^i$ by setting an appropriate schedule $a^i(\epsilon)$ for purchasing state-contingent -securities. +Consumer of type $i$ can offset any choice of $\theta^i$ by setting an appropriate schedule $a^i(\epsilon)$ for purchasing state-contingent securities. ### Computing competitive equilibrium prices and quantities -Having computed an allocation that solves the planning problem, we can -readily compute a competitive equilibrium via the following steps that, -as we’ll see, relies heavily on the `Big K, little k`, -`Big C, little c` logic mentioned earlier: +Having computed an allocation that solves the planning problem, we can readily compute a competitive equilibrium via the following steps that, as we’ll see, rely heavily on the `Big K, little k`, `Big C, little c` logic mentioned earlier: -- a competitive equilbrium allocation equals the allocation chosen by +- a competitive equilibrium allocation equals the allocation chosen by the planner - competitive equilibrium prices and the value of a firm’s equity are encoded in shadow prices from the planning problem that depend on Big $K$ and Big $C$. To substantiate that this procedure is valid, we proceed as follows. -With $K$ in hand, we make the following guess for competitive -equilibrium Arrow securities prices +With $K$ in hand, we make the following guess for competitive equilibrium Arrow securities prices ```{math} :label: arrowprices -q(\epsilon;K) = \beta \left( \frac{u'\left( w_1(\epsilon) + A K^\alpha e^\epsilon\right)} {u'(w_0 - K )} \right)^{-\gamma} +q(\epsilon;K) = \beta \frac{u'\left( w_1(\epsilon) + A K^\alpha e^\epsilon\right)} {u'(w_0 - K )} g(\epsilon) += \beta \left( \frac{w_1(\epsilon) + A K^\alpha e^\epsilon}{w_0 - K} \right)^{-\gamma} g(\epsilon) ``` -To confirm the guess, we begin by considering its consequences for the firm’s choice of $k$. +To confirm the guess, we begin by considering its consequences for the firm’s choice of $k$. -With Arrow securities prices {eq}`arrowprices`, the firm’s -first-order necessary condition for choosing $k$ becomes +With Arrow securities prices {eq}`arrowprices`, the firm’s first-order necessary condition for choosing $k$ becomes ```{math} :label: kK @@ -660,13 +573,9 @@ $$ k = K $$ -because by setting $k=K$ equation {eq}`kK` becomes -equivalent with the planner’s first-order condition -{eq}`focke` for setting $K$. +because by setting $k=K$ equation {eq}`kK` becomes equivalent with the planner’s first-order condition {eq}`focke` for setting $K$. -To pose a consumer’s problem in a competitive equilibrium, we require -not only the above guess for the Arrow securities pricing kernel -$q(\epsilon)$ but the value of equity $\tilde V$: +To pose a consumer’s problem in a competitive equilibrium, we require not only the above guess for the Arrow securities pricing kernel $q(\epsilon)$ but also the value of equity $\tilde V$: ```{math} :label: tildeV2 @@ -674,40 +583,32 @@ $q(\epsilon)$ but the value of equity $\tilde V$: \tilde V = \int A K^\alpha e^\epsilon q(\epsilon;K) d \epsilon ``` -Let $\tilde V$ be the value of equity implied by Arrow securities -price function {eq}`arrowprices` and formula -{eq}`tildeV2`. +Let $\tilde V$ be the value of equity implied by Arrow securities price function {eq}`arrowprices` and formula {eq}`tildeV2`. -At the Arrow securities prices $q(\epsilon)$ given by {eq}`arrowprices` -and equity value $\tilde V$ given by {eq}`tildeV2`, -consumer $i=1,2$ choose consumption allocations and portolios -that satisfy the first-order necessary conditions +At the Arrow securities prices $q(\epsilon)$ given by {eq}`arrowprices` and equity value $\tilde V$ given by {eq}`tildeV2`, consumers $i=1,2$ choose consumption allocations and portfolios that satisfy the first-order necessary conditions $$ \beta \left( \frac{u'(c_1^i(\epsilon))}{u'(c_0^i)} \right) g(\epsilon) = q(\epsilon;K) $$ -It can be verified directly that the following choices satisfy these -equations +It can be verified directly that the following choices satisfy these equations $$ \begin{aligned} c_0^1 + c_0^2 & = C_0 = w_0 - K \cr -c_0^1(\epsilon) + c_0^2(\epsilon) & = C_1(\epsilon) = w_1(\epsilon) + A k^\alpha e ^\epsilon \cr +c_1^1(\epsilon) + c_1^2(\epsilon) & = C_1(\epsilon) = w_1(\epsilon) + A K^\alpha e^\epsilon \cr \frac{c_1^2(\epsilon)}{c_1^1(\epsilon)} & = \frac{c_0^2}{c_0^1} = \frac{1-\eta}{\eta} \end{aligned} $$ -for an $\eta \in (0,1)$ that depends on consumers’ -endowments -$[w_0^1, w_0^2, w_1^1(\epsilon), w_1^2(\epsilon), \theta_0^1, \theta_0^2 ]$. +for an $\eta \in (0,1)$ that depends on consumers’ endowments $[w_0^1, w_0^2, w_1^1(\epsilon), w_1^2(\epsilon), \theta_0^1, \theta_0^2 ]$. + +**Remark:** Multiple arrangements of endowments $[w_0^1, w_0^2, w_1^1(\epsilon), w_1^2(\epsilon), \theta_0^1, \theta_0^2 ]$ are associated with the same distribution of wealth $\eta$. -**Remark:** Multiple arrangements of endowments -$[w_0^1, w_0^2, w_1^1(\epsilon), w_1^2(\epsilon), \theta_0^1, \theta_0^2 ]$ -associated with the same distribution of wealth $\eta$. Can you explain why? +Can you explain why? ```{hint} -Think about the portfolio indeterminacy finding above. +Consumer $i$'s budget constraint {eq}`noarb` depends on the endowments $w_0^i, \theta_0^i, w_1^i(\epsilon)$ only through their present value $w_0^i + \theta_0^i V + \int w_1^i(\epsilon) q(\epsilon) d\epsilon$. ``` ### Modigliani-Miller theorem @@ -723,29 +624,19 @@ d^b(k,b;\epsilon) &= \min \left\{ \frac{e^\epsilon A k^\alpha}{b}, 1 \right\} \end{aligned} $$ -Thus, one unit of the bond pays one unit of consumption at time -$1$ in state $\epsilon$ if -$A k^\alpha e^\epsilon - b \geq 0$, which is true when -$\epsilon \geq \epsilon^* = \log \frac{b}{Ak^\alpha}$, and pays -$\frac{A k^\alpha e^\epsilon}{b}$ units of time $1$ -consumption in state $\epsilon$ when -$\epsilon < \epsilon^*$. +Thus, one unit of the bond pays one unit of consumption at time $1$ in state $\epsilon$ if $A k^\alpha e^\epsilon - b \geq 0$, which is true when $\epsilon \geq \epsilon^* = \log \frac{b}{Ak^\alpha}$, and pays $\frac{A k^\alpha e^\epsilon}{b}$ units of time $1$ consumption in state $\epsilon$ when $\epsilon < \epsilon^*$. -The value of the firm is now the sum of equity plus the value of bonds, -which we denote +The market value of the firm's securities is now the sum of the value of its equity and the value of its bonds, which we denote $$ \tilde V + b p(k,b) $$ -where $p(k,b)$ is the price of one unit of the bond when a firm -with $k$ units of physical capital issues $b$ bonds. +where $p(k,b)$ is the price of one unit of the bond when a firm with $k$ units of physical capital issues $b$ bonds. -We continue to assume that there are complete markets in Arrow -securities with pricing kernel $q(\epsilon)$. +We continue to assume that there are complete markets in Arrow securities with pricing kernel $q(\epsilon)$. -A version of the no-arbitrage-in-equilibrium argument that we presented -earlier implies that the value of equity and the price of bonds are +A version of the no-arbitrage-in-equilibrium argument that we presented earlier implies that the value of equity and the price of bonds are $$ \begin{aligned} @@ -755,67 +646,69 @@ p(k, b) & = \frac{A k^\alpha}{b} \int_{-\infty}^{\epsilon^*} e^\epsilon q(\eps \end{aligned} $$ -Consequently, the value of the firm is +Consequently, the market value of the firm's securities is $$ \tilde V + p(k,b) b = A k^\alpha \int_{-\infty}^\infty e^\epsilon q(\epsilon) d \epsilon, $$ -which is the same expression that we obtained above when we assumed that -the firm issued only equity. +which is the same expression that we obtained above when we assumed that the firm issued only equity. -We thus obtain a version of the celebrated Modigliani-Miller theorem {cite}`Modigliani_Miller_1958` -about firms’ finance: +Owners still choose $k$ to maximize the value of the firm net of its capital outlay, $V = -k + \tilde V + p(k,b) b$. + +Because $\tilde V + p(k,b) b$ does not depend on $b$, the first-order condition for $k$ is again {eq}`kK`. + +We thus obtain a version of the celebrated Modigliani-Miller theorem {cite}`Modigliani_Miller_1958` about firms’ finance: **Modigliani-Miller theorem:** -- The value of a firm is independent the mix of equity and bonds that +- The value of a firm is independent of the mix of equity and bonds that it uses to finance its physical capital. -- The firms’s decision about how much physical capital to purchase does +- The firm’s decision about how much physical capital to purchase does not depend on whether it finances those purchases by issuing bonds - or equity + or equity. - The firm’s choice of whether to finance itself by issuing equity or - bonds is indeterminant + bonds is indeterminate. -Please note the role of the assumption of complete markets in Arrow -securities in substantiating these claims. +Please note the role of the assumption of complete markets in Arrow securities in substantiating these claims. -In {doc}`Equilibrium Capital Structures with Incomplete Markets `, we will assume that markets are (very) -incomplete – we’ll shut down markets in almost all Arrow securities. +In {doc}`Equilibrium Capital Structures with Incomplete Markets `, we will assume that markets are (very) incomplete – we’ll shut down markets in almost all Arrow securities. That will pull the rug from underneath the Modigliani-Miller theorem. ## Code -We create a class object `BCG_complete_markets` to compute -equilibrium allocations of the complete market BCG model given a list -of parameter values. +We create a class object `BCG_complete_markets` to compute equilibrium allocations of the complete market BCG model given a list of parameter values. It consists of 4 functions that do the following things: * `opt_k` computes the planner's optimal capital $K$ - - First, create a grid for capital. - - Then for each value of capital stock in the grid, compute the left side of the planner's - first-order necessary condition for $k$, that is, + - It uses a Newton-secant root finder, started at $k = 0.01$, to find the $k$ that solves the planner's + first-order necessary condition {eq}`focke`, written as $$ - \beta \alpha A K^{\alpha -1} \int \left( \frac{w_1(\epsilon) + A K^\alpha e^\epsilon}{w_0 - K } \right)^{-\gamma} e^\epsilon g(\epsilon) d \epsilon - 1 =0 + \beta \alpha A k^{\alpha -1} \int \left( \frac{w_1(\epsilon) + A k^\alpha e^\epsilon}{w_0 - k } \right)^{-\gamma} e^\epsilon g(\epsilon) d \epsilon - 1 = 0 $$ - - Find $k$ that solves this equation. -* `q` computes Arrow security prices as a function of the productivity shock $\epsilon$ and capital $K$: + where the integral is computed by Gauss-Hermite quadrature. + - When called with `plot=True`, it also plots the left side of this equation on a grid of values of $k$. +* `q` computes the ratio of Arrow security prices to probabilities as a function of the productivity shock $\epsilon$ and capital $K$: $$ - q(\epsilon;K) = \beta \left( \frac{u'\left( w_1(\epsilon) + A K^\alpha e^\epsilon\right)} {u'(w_0 - K )} \right) + \frac{q(\epsilon;K)}{g(\epsilon)} = \beta \frac{u'\left( w_1(\epsilon) + A K^\alpha e^\epsilon\right)} {u'(w_0 - K )} $$ -* `V` solves for the firm value given capital $k$: + so the Arrow price $q(\epsilon;K)$ in {eq}`arrowprices` equals `q` times the density $g(\epsilon)$. + +* `V` solves for the firm value given capital $k$, evaluating Arrow prices at $K = k$: $$ - V = - k + \int A k^\alpha e^\epsilon q(\epsilon; K) d \epsilon + V = - k + \int A k^\alpha e^\epsilon q(\epsilon; K) d \epsilon = - k + \int A k^\alpha e^\epsilon \frac{q(\epsilon; K)}{g(\epsilon)} g(\epsilon) d \epsilon $$ -* `opt_c` computes optimal consumptions $c^i_0$, and $c^i(\epsilon)$: + where the second integral, an expectation with respect to $g$, is computed by Gauss-Hermite quadrature. + +* `opt_c` computes optimal consumptions $c^i_0$ and $c^i_1(\epsilon)$: - The function first computes weight $\eta$ using the budget constraint for agent 1: @@ -824,6 +717,7 @@ It consists of 4 functions that do the following things: = c_0^1 + \int c_1^1(\epsilon) q(\epsilon) d \epsilon = \eta \left( C_0 + \int C_1(\epsilon) q(\epsilon) d \epsilon \right) $$ + where $$ @@ -844,21 +738,20 @@ It consists of 4 functions that do the following things: \end{aligned} $$ - The list of parameters includes: -- $\chi_1$, $\chi_2$: Correlation parameters for agents 1 - and 2. Default values are 0 and 0.9, respectively. +- $\chi_1$, $\chi_2$: loadings of the log endowments of agents 1 + and 2 on the shock $\epsilon$. Default values are 0 and 0.9, respectively. - $w^1_0$, $w^2_0$: Initial endowments. Default values are 1. - $\theta^1_0$, $\theta^2_0$: Consumers’ initial shares of a representative firm. Default values are 0.5. -- $\psi$: CRRA risk parameter. Default value is 3. -- $\alpha$: Returns to scale production function parameter. +- $\psi$: CRRA risk parameter, denoted $\gamma$ in the text above. Default value is 3. +- $\alpha$: Capital share (curvature) parameter of the production function. Default value is 0.6. - $A$: Productivity of technology. Default value is 2.5. - $\mu$, $\sigma$: Mean and standard deviation of the log of the shock. Default values are -0.025 and 0.4, respectively. -- $\beta$: time preference discount factor. Default value is .96. +- $\beta$: time preference discount factor. Default value is 0.96. - `nb_points_integ`: number of points used for integration through Gauss-Hermite quadrature: default value is 10 @@ -895,7 +788,7 @@ class BCG_complete_markets: self.𝜒1 = 𝜒1 self.𝜒2 = 𝜒2 - # Other parameters + # Other parameters (𝜓 is the CRRA parameter 𝛾 in the text) self.𝜓 = 𝜓 self.𝛼 = 𝛼 self.A = A @@ -903,9 +796,6 @@ class BCG_complete_markets: self.𝜎 = 𝜎 self.𝛽 = 𝛽 - # Utility - self.u = lambda c: (c**(1-𝜓)) / (1-𝜓) - # Production self.f = njit(lambda k: A * (k ** 𝛼)) self.Y = lambda 𝜖, k: np.exp(𝜖) * self.f(k) @@ -941,17 +831,11 @@ class BCG_complete_markets: def opt_k(self, plot=False): w0 = self.w0 - # Grid for k - kgrid = np.linspace(1e-4, w0-1e-4, 100) - - # get FONC values for each k in the grid - kfoc_list = []; - for k in kgrid: - kfoc = self.k_foc(k, self.𝜒1, self.𝜒2) - kfoc_list.append(kfoc) - - # Plot FONC for k + # Plot FONC for k on a grid if plot: + kgrid = np.linspace(1e-4, w0-1e-4, 100) + kfoc_list = [self.k_foc(k, self.𝜒1, self.𝜒2) for k in kgrid] + fig, ax = plt.subplots(figsize=(8,7)) ax.plot(kgrid, kfoc_list, color='blue', label=r'FONC for k') ax.axhline(0, color='red', linestyle='--') @@ -972,8 +856,8 @@ class BCG_complete_markets: w0 = self.w0 w1 = self.w1 fk = self.f(k) - g = self.g + # Returns 𝛽 u'(C_1(𝜖))/u'(C_0), i.e., the Arrow price q(𝜖;K) divided by g(𝜖) return 𝛽 * ((w1(𝜖) + np.exp(𝜖)*fk) / (w0 - k))**(-𝜓) @@ -1029,7 +913,6 @@ def k_foc_factory(model): 𝛽 = model.𝛽 𝛼 = model.𝛼 A = model.A - 𝜓 = model.𝜓 w0 = model.w0 𝜇 = model.𝜇 𝜎 = model.𝜎 @@ -1037,7 +920,7 @@ def k_foc_factory(model): weights = model.weights points_integral = model.points_integral - w11 = njit(lambda 𝜖, 𝜒1, : np.exp(-𝜒1*𝜇 - 0.5*(𝜒1**2)*(𝜎**2) + 𝜒1*𝜖)) + w11 = njit(lambda 𝜖, 𝜒1: np.exp(-𝜒1*𝜇 - 0.5*(𝜒1**2)*(𝜎**2) + 𝜒1*𝜖)) w21 = njit(lambda 𝜖, 𝜒2: np.exp(-𝜒2*𝜇 - 0.5*(𝜒2**2)*(𝜎**2) + 𝜒2*𝜖)) w1 = njit(lambda 𝜖, 𝜒1, 𝜒2: w11(𝜖, 𝜒1) + w21(𝜖, 𝜒2)) @@ -1060,19 +943,15 @@ def k_foc_factory(model): ### Examples -Below we provide some examples of how to use `BCG_complete markets`. +Below we provide some examples of how to use `BCG_complete_markets`. -#### 1st example +#### First example -In the first example, we set up instances of BCG complete markets -models. +In the first example, we set up instances of BCG complete markets models. -We can use either default parameter values or set parameter values as we -want. +We can use either default parameter values or set parameter values as we want. -The two instances of the BCG complete markets model, `mdl1` and -`mdl2`, represent the model with default parameter settings and with agent 2’s income correlation altered to be $\chi_2 = -0.9$, -respectively. +The two instances of the BCG complete markets model, `mdl1` and `mdl2`, represent the model with default parameter settings and with the loading of agent 2’s endowment on the shock altered to be $\chi_2 = -0.9$, respectively. ```{code-cell} python3 # Example: BCG model for complete markets @@ -1080,18 +959,17 @@ mdl1 = BCG_complete_markets() mdl2 = BCG_complete_markets(𝜒2=-0.9) ``` -Let’s plot the agents’ time-1 endowments with respect to shocks to see -the difference in the two models: +Let’s plot the agents’ time-1 endowments as functions of the shock to see how the two models differ. ```{code-cell} python3 #==== Figure 1: HH endowments and firm productivity ====# -# Realizations of innovation from -3 to 3 +# Realizations of innovation from -1 to 1 epsgrid = np.linspace(-1,1,1000) fig, ax = plt.subplots(1,2,figsize=(14,6)) -ax[0].plot(epsgrid, mdl1.w11(epsgrid), color='black', label=r'Agent 1\'s endowment') -ax[0].plot(epsgrid, mdl1.w21(epsgrid), color='blue', label=r'Agent 2\'s endowment') +ax[0].plot(epsgrid, mdl1.w11(epsgrid), color='black', label="Agent 1's endowment") +ax[0].plot(epsgrid, mdl1.w21(epsgrid), color='blue', label="Agent 2's endowment") ax[0].plot(epsgrid, mdl1.Y(epsgrid,1), color='red', label=r'Production with $k=1$') ax[0].set_xlim([-1,1]) ax[0].set_ylim([0,7]) @@ -1100,8 +978,8 @@ ax[0].set_title(r'Model with $\chi_1 = 0$, $\chi_2 = 0.9$') ax[0].legend() ax[0].grid() -ax[1].plot(epsgrid, mdl2.w11(epsgrid), color='black', label=r'Agent 1\'s endowment') -ax[1].plot(epsgrid, mdl2.w21(epsgrid), color='blue', label=r'Agent 2\'s endowment') +ax[1].plot(epsgrid, mdl2.w11(epsgrid), color='black', label="Agent 1's endowment") +ax[1].plot(epsgrid, mdl2.w21(epsgrid), color='blue', label="Agent 2's endowment") ax[1].plot(epsgrid, mdl2.Y(epsgrid,1), color='red', label=r'Production with $k=1$') ax[1].set_xlim([-1,1]) ax[1].set_ylim([0,7]) @@ -1113,8 +991,7 @@ ax[1].grid() plt.show() ``` -Let’s also compare the optimal capital stock, $k$, and optimal -time-0 consumption of agent 2, $c^2_0$, for the two models: +Let’s also compare the optimal capital stock, $k$, and optimal time-0 consumption of agent 2, $c^2_0$, for the two models: ```{code-cell} python3 # Print optimal k @@ -1132,13 +1009,13 @@ print('The optimal c20 for model 1: {:.5f}'.format(c20_1)) print('The optimal c20 for model 2: {:.5f}'.format(c20_2)) ``` -#### 2nd example +#### Second example + +In the second example, we illustrate how the optimal choice of $k$ is influenced by the loadings $\chi_i$ of endowments on the shock. -In the second example, we illustrate how the optimal choice of $k$ -is influenced by the correlation parameter $\chi_i$. +We will need to install the `plotly` package for 3D illustration. -We will need to install the `plotly` package for 3D illustration. See -[https://plotly.com/python/getting-started/](https://plotly.com/python/getting-started/) for further instructions. +See [https://plotly.com/python/getting-started/](https://plotly.com/python/getting-started/) for further instructions. ```{code-cell} python3 # Mesh grid of 𝜒 @@ -1151,8 +1028,6 @@ k_foc = k_foc_factory(mdl1) # Create grid for k kgrid = np.zeros_like(𝜒1grid) -w0 = mdl1.w0 - @njit(parallel=True) def fill_k_grid(kgrid): # Loop: Compute optimal k and @@ -1176,7 +1051,7 @@ fill_k_grid(kgrid) ``` ```{code-cell} python3 -#=== Example: Plot optimal k with different correlations ===# +#=== Example: Plot optimal k with different loadings ===# from IPython.display import Image # Import plotly @@ -1198,3 +1073,302 @@ Image(fig.to_image(format="png", engine="kaleido")) # fig.show() will provide interactive plot when running # notebook locally ``` + +The optimal $k$ is smallest, about $0.129$, near $\chi_1 = \chi_2 = 0$, where second-period endowments are riskless. + +It rises as endowments load on the shock in either direction, and it is largest, about $0.201$, at $\chi_1 = \chi_2 = 1$, where endowment risk reinforces productivity risk. + +When the loadings have opposite signs, as at $\chi_1 = -\chi_2 = \pm 1$, the two endowments partly offset one another and $k$ is only about $0.134$. + +This pattern is consistent with a precautionary motive: riskier second-period endowments raise the expected marginal utility of second-period consumption and with it the incentive to carry resources into period $1$ by accumulating capital. + +## Exercises + +```{exercise} +:label: bcgc_ex1 + +This exercise verifies the Modigliani-Miller theorem numerically at the equilibrium computed by `BCG_complete_markets`. + +Note that the method `q(𝜖, k)` of `BCG_complete_markets` returns $\beta u'(C_1(\epsilon))/u'(C_0)$, so the Arrow price density that appears in the formulas for $\tilde V$ and $p(k,b)$ is `q(𝜖, K) * g(𝜖)`. + +1. Using the default parameters, compute $K$ and, for $b \in \{0.01, 0.2, 0.5, 0.8, 1.2, 2.0\}$, compute the default threshold $\epsilon^*$, the value of equity $\tilde V$, the bond price $p(K,b)$, the value of debt $p(K,b) b$, and the market value of the firm's securities $\tilde V + p(K,b) b$. + +2. Holding Arrow prices fixed at $q(\epsilon;K)$, let a firm that has promised to issue $b$ bonds choose $k$ to maximize $-k + \tilde V(k,b) + p(k,b) b$. Show that its optimal $k$ does not depend on $b$. + +Hint: because the payoffs $d^e$ and $d^b$ have a kink at $\epsilon^*$, integrate with `scipy.integrate.quad` on either side of $\epsilon^*$ rather than with the lecture's Gauss-Hermite nodes. +``` + +```{solution-start} bcgc_ex1 +:class: dropdown +``` + +Here is one solution. + +```{code-cell} ipython3 +from scipy.integrate import quad +from scipy.optimize import minimize_scalar + +mdl = BCG_complete_markets() +K = mdl.opt_k() +lo, hi = mdl.𝜇 - 10 * mdl.𝜎, mdl.𝜇 + 10 * mdl.𝜎 + +def arrow_price(𝜖, K): + "Arrow price density q(𝜖;K) = 𝛽 u'(C_1)/u'(C_0) g(𝜖)" + return mdl.q(𝜖, K) * mdl.g(𝜖) + +def equity_bond_values(k, b, K): + "Value of equity and price of a bond for a firm with (k, b), at prices q(.;K)" + Y = mdl.f(k) + 𝜖_star = np.log(b / Y) + Vtilde = quad(lambda 𝜖: (Y * np.exp(𝜖) - b) * arrow_price(𝜖, K), + max(𝜖_star, lo), hi)[0] + p = (quad(lambda 𝜖: Y * np.exp(𝜖) / b * arrow_price(𝜖, K), + lo, min(𝜖_star, hi))[0] + + quad(lambda 𝜖: arrow_price(𝜖, K), max(𝜖_star, lo), hi)[0]) + return Vtilde, p + +print(f"K = {K:.5f}, A K^alpha = {mdl.f(K):.4f}") +print(f"{'b':>6} {'eps*':>8} {'Vtilde':>9} {'p':>8} {'p*b':>8} {'Vtilde+p*b':>11}") +for b in [0.01, 0.2, 0.5, 0.8, 1.2, 2.0]: + Vt, p = equity_bond_values(K, b, K) + print(f"{b:6.2f} {np.log(b/mdl.f(K)):8.3f} {Vt:9.5f} {p:8.5f} {p*b:8.5f} {Vt + p*b:11.6f}") + +print(f"\nPrice of a riskless claim, int q: {quad(lambda e: arrow_price(e, K), lo, hi)[0]:.5f}") + +# Optimal k for different leverage levels, prices held fixed at q(.;K) +for b in [0.01, 0.5, 1.2]: + obj = lambda k: -(-k + sum(np.array(equity_bond_values(k, b, K)) * np.array([1, b]))) + res = minimize_scalar(obj, bounds=(0.02, 0.6), method='bounded', + options={'xatol': 1e-7}) + print(f"b = {b:4.2f}: argmax_k [-k + Vtilde + p b] = {res.x:.5f}") +``` + +Several things stand out. + +As $b$ rises from $0.01$ to $2$, the default threshold $\epsilon^*$ rises from about $-4.35$ to about $0.95$, so default goes from essentially impossible to more likely than not. + +The value of equity falls from $0.2335$ to almost zero and the bond price falls from $0.3771$ to $0.1186$, but the market value of the firm's securities $\tilde V + p b$ equals $0.237247$ for every $b$. + +That number equals $K + V$ from the lecture's `V` method: $0.14235 + 0.09490$. + +When $b$ is small the bond is effectively riskless, so its price equals the price $\int q(\epsilon;K) d\epsilon = 0.3771$ of a sure unit of time $1$ consumption. + +The firm's optimal $k$ is $0.14235 = K$ whatever $b$ is. + +The economics is that with complete Arrow markets, equity and debt are simply two bundles of Arrow securities whose payoffs add up to $A k^\alpha e^\epsilon$ in every state. + +Because Arrow securities are priced linearly, how the firm slices its output between the two bundles cannot change the value of the total, and therefore cannot change the $k$ that maximizes $-k$ plus that value. + +```{solution-end} +``` + +```{exercise} +:label: bcgc_ex2 + +This exercise studies how the planner's (and the competitive equilibrium's) capital $K$ responds to risk aversion and to risk. + +Define the stochastic discount factor $m(\epsilon) = \beta u'(C_1(\epsilon))/u'(C_0)$, the gross riskless rate $R_f = 1/E[m]$, and the marginal gross return on capital $R_k(\epsilon) = \alpha A K^{\alpha-1} e^\epsilon$. + +Condition {eq}`focke` says $E[m R_k] = 1$, which implies $E[R_k] - R_f = -R_f \, \textrm{cov}(m, R_k)$. + +1. For $\psi \in \{1.5, 2, 3, 4, 6\}$, holding other parameters at their defaults, compute $K$, $R_f$, $E[R_k]$, $E[R_k]-R_f$ and $\textrm{cov}(m,R_k)$. + +2. Do the same for $\sigma \in \{0.05, 0.2, 0.4, 0.6, 0.8\}$, setting $\mu = -\sigma^2/2$ so that $E[e^\epsilon] = 1$ and each increase in $\sigma$ is a mean-preserving spread in $e^\epsilon$ (and in each $w_1^i(\epsilon)$). + +3. Do the same for $\chi_2 \in \{-0.9, 0, 0.9\}$ at the default parameters. + +4. Plot $K$ as a function of $\psi$ and of $\sigma$ and explain what you find. +``` + +```{solution-start} bcgc_ex2 +:class: dropdown +``` + +Here is one solution. + +```{code-cell} ipython3 +def expect(mdl, h): + "E[h(𝜖)] under N(𝜇, 𝜎^2) by Gauss-Hermite quadrature" + return mdl.weights @ h(mdl.points_integral) / np.sqrt(np.pi) + +def summarize(mdl): + K = mdl.opt_k() + m = lambda 𝜖: mdl.q(𝜖, K) # stochastic discount factor + Rf = 1 / expect(mdl, m) # gross riskless rate + MPK = lambda 𝜖: mdl.𝛼 * mdl.f(K) / K * np.exp(𝜖) # marginal return on capital + ER = expect(mdl, MPK) # expected marginal return + cov = expect(mdl, lambda 𝜖: m(𝜖) * MPK(𝜖)) - expect(mdl, m) * ER + return K, Rf, ER, cov + +print("Varying 𝜓 (𝜎 = 0.4)") +print(f"{'𝜓':>5} {'K':>8} {'Rf':>7} {'E[R_k]':>8} {'E[R_k]-Rf':>10} {'cov(m,R_k)':>11}") +for 𝜓 in [1.5, 2, 3, 4, 6]: + K, Rf, ER, cov = summarize(BCG_complete_markets(𝜓=𝜓, nb_points_integ=20)) + print(f"{𝜓:5.1f} {K:8.5f} {Rf:7.4f} {ER:8.4f} {ER-Rf:10.4f} {cov:11.5f}") + +print("\nMean-preserving spreads: 𝜇 = -𝜎^2/2 so that E[exp(𝜖)] = 1 (𝜓 = 3)") +print(f"{'𝜎':>5} {'K':>8} {'Rf':>7} {'E[R_k]':>8} {'E[R_k]-Rf':>10} {'cov(m,R_k)':>11}") +for 𝜎 in [0.05, 0.2, 0.4, 0.6, 0.8]: + K, Rf, ER, cov = summarize(BCG_complete_markets(𝜎=𝜎, 𝜇=-𝜎**2/2, nb_points_integ=20)) + print(f"{𝜎:5.2f} {K:8.5f} {Rf:7.4f} {ER:8.4f} {ER-Rf:10.4f} {cov:11.5f}") + +print("\nVarying 𝜒2 (𝜓 = 3, 𝜎 = 0.4, 𝜇 = -0.025)") +for 𝜒2 in [-0.9, 0, 0.9]: + K, Rf, ER, cov = summarize(BCG_complete_markets(𝜒2=𝜒2, nb_points_integ=20)) + print(f"{𝜒2:5.1f} {K:8.5f} {Rf:7.4f} {ER:8.4f} {ER-Rf:10.4f} {cov:11.5f}") + +𝜓_grid = np.linspace(1.2, 8, 15) +K_𝜓 = [BCG_complete_markets(𝜓=𝜓).opt_k() for 𝜓 in 𝜓_grid] +𝜎_grid = np.linspace(0.05, 0.8, 15) +K_𝜎 = [BCG_complete_markets(𝜎=𝜎, 𝜇=-𝜎**2/2).opt_k() for 𝜎 in 𝜎_grid] + +fig, ax = plt.subplots(1, 2, figsize=(12, 4.5)) +ax[0].plot(𝜓_grid, K_𝜓) +ax[0].set_xlabel(r'$\psi$') +ax[0].set_ylabel(r'$K$') +ax[1].plot(𝜎_grid, K_𝜎) +ax[1].set_xlabel(r'$\sigma$ (with $\mu = -\sigma^2/2$)') +ax[1].set_ylabel(r'$K$') +plt.tight_layout() +plt.show() +print("K strictly decreasing in 𝜓:", np.all(np.diff(K_𝜓) < 0)) +print("K strictly increasing in 𝜎:", np.all(np.diff(K_𝜎) > 0)) +``` + +In every row, $E[R_k] - R_f = -R_f \, \textrm{cov}(m, R_k)$, as it must: for example, at the default parameters $2.6518 \times 0.30345 = 0.8047$. + +**Risk aversion.** + +$K$ falls steadily as $\psi$ rises, from $0.267$ at $\psi = 1.5$ to $0.086$ at $\psi = 6$. + +In this economy consumption is expected to grow ($C_1$ is on average well above $C_0 \approx 1.86$), so a larger $\psi$, which in this time-separable specification also means a smaller elasticity of intertemporal substitution, makes consumers less willing to postpone consumption. + +At the same time, the risk premium on capital rises from $0.37$ to $1.56$ because capital pays off most in states where consumption is high. + +Both forces lower $K$. + +$R_f$ is not monotone in $\psi$: it rises from $2.31$ to $2.72$ and then falls back to $2.67$ at $\psi = 6$, as the precautionary motive, which grows with $\psi$, starts to offset the consumption-smoothing motive. + +**Risk.** + +With mean-preserving spreads, $K$ rises with $\sigma$, from $0.1349$ at $\sigma = 0.05$ to $0.1498$ at $\sigma = 0.8$. + +A larger $\sigma$ raises the risk premium from $0.015$ to $1.81$, which by itself discourages investment. + +But it also strengthens the precautionary saving motive so much that $R_f$ falls from $3.33$ to $1.40$. + +Since capital is the only way to move goods from period $0$ to period $1$ in the aggregate, the precautionary motive wins and $E[R_k]$ (equal to $\alpha A K^{\alpha-1} E[e^\epsilon]$) must fall from $3.34$ to $3.21$, which requires more capital. + +**Loadings of endowments on the productivity shock.** + +$K$ is lowest at $\chi_2 = 0$ ($0.1292$), when the aggregate endowment $w_1(\epsilon) = 2$ is riskless. + +With $\chi_2 = -0.9$, agent 2's endowment hedges output, so capital is almost riskless in terms of marginal utility: the risk premium is only $0.014$ and, although $R_f = 3.49$ is higher than at $\chi_2 = 0$, the required expected return $E[R_k] = 3.50$ is lower than at $\chi_2 = 0$ ($3.59$), so $K = 0.1379$ is higher. + +With $\chi_2 = 0.9$, the endowment amplifies aggregate risk: the risk premium is largest ($0.80$) but $R_f$ is lowest ($2.65$), and $K = 0.1424$ is highest. + +```{solution-end} +``` + +```{exercise} +:label: bcgc_ex3 + +The lecture remarks that multiple arrangements of endowments $[w_0^1, w_0^2, w_1^1(\epsilon), w_1^2(\epsilon), \theta_0^1, \theta_0^2]$ are associated with the same distribution of wealth $\eta$. + +1. Verify this numerically: starting from the default parameters, reassign all shares of the firm to agent 1 (and, separately, to agent 2), and offset the change by transferring time $0$ endowment so that $w_0^i + \theta_0^i V$ is unchanged for each $i$. +Report $K$, $\eta$, and the Pareto weight $\phi_1$ that the planner would need to attach to agent 1 to implement the same allocation. +Also report what happens if all shares go to agent 1 without an offsetting transfer. + +2. Illustrate portfolio indeterminacy: at the default equilibrium, for $\theta^1 \in \{0, 0.5, 1, -1\}$ compute the Arrow security holdings $a^1(\epsilon)$ that deliver agent 1's equilibrium consumption $c_1^1(\epsilon)$, and show that the time $0$ cost $\int q(\epsilon) a^1(\epsilon) d\epsilon + \theta^1 \tilde V$ of the portfolio does not depend on $\theta^1$. + +3. Plot $\beta u'(C_1(\epsilon))/u'(C_0)$ and the Arrow price density $q(\epsilon;K)$ for $\chi_2 = 0.9$ and $\chi_2 = -0.9$, and explain their shapes. +``` + +```{solution-start} bcgc_ex3 +:class: dropdown +``` + +Here is one solution. + +```{code-cell} ipython3 +def expect(mdl, h): + return mdl.weights @ h(mdl.points_integral) / np.sqrt(np.pi) + +base = BCG_complete_markets() +K = base.opt_k() +V = base.V(K) +print(f"K = {K:.5f}, V = {V:.5f}") + +# Part 1: endowment arrangements with the same distribution of wealth +arrangements = { + "baseline (θ10=0.5, w10=1)": dict(𝜃10=0.5, 𝜃20=0.5, w10=1.0, w20=1.0), + "agent 1 owns the firm": dict(𝜃10=1.0, 𝜃20=0.0, w10=1-0.5*V, w20=1+0.5*V), + "agent 2 owns the firm": dict(𝜃10=0.0, 𝜃20=1.0, w10=1+0.5*V, w20=1-0.5*V), + "agent 1 owns firm, no offset": dict(𝜃10=1.0, 𝜃20=0.0, w10=1.0, w20=1.0), +} +𝜓 = base.𝜓 +for name, kw in arrangements.items(): + mdl = BCG_complete_markets(**kw) + k = mdl.opt_k() + c10, c20, c11, c21 = mdl.opt_c(k=k) + 𝜂 = c10 / (mdl.w0 - k) + 𝜙1 = 𝜂**𝜓 / (𝜂**𝜓 + (1 - 𝜂)**𝜓) + print(f"{name:32s} K = {k:.5f} eta = {𝜂:.5f} phi_1 = {𝜙1:.5f}") + +# Part 2: two portfolios that finance agent 1's same consumption plan +c10, c20, c11, c21 = base.opt_c(k=K) +m = lambda 𝜖: base.q(𝜖, K) +Vtilde = V + K +for 𝜃 in [0.0, 0.5, 1.0, -1.0]: + a = lambda 𝜖: c11(𝜖) - base.w11(𝜖) - 𝜃 * base.Y(𝜖, K) # Arrow purchases + cost = expect(base, lambda 𝜖: m(𝜖) * a(𝜖)) + 𝜃 * Vtilde + print(f"theta1 = {𝜃:5.2f}: time-0 cost of portfolio = {cost:.6f}, " + f"implied c10 = {base.w10 + base.𝜃10*V - cost:.6f}") + +# Part 3: Arrow price densities for chi2 = 0.9 and -0.9 +epsgrid = np.linspace(-1.5, 1.5, 400) +fig, ax = plt.subplots(1, 2, figsize=(12, 4.5)) +for chi2, ls in [(0.9, '-'), (-0.9, '--')]: + mdl = BCG_complete_markets(𝜒2=chi2) + Kc = mdl.opt_k() + ax[0].plot(epsgrid, mdl.q(epsgrid, Kc), ls, label=rf'$\chi_2 = {chi2}$') + ax[1].plot(epsgrid, mdl.q(epsgrid, Kc) * mdl.g(epsgrid), ls, label=rf'$\chi_2 = {chi2}$') + print(f"chi2 = {chi2:4.1f}: m(-1) = {mdl.q(-1.0, Kc):.4f}, m(0) = {mdl.q(0.0, Kc):.4f}, " + f"m(1) = {mdl.q(1.0, Kc):.4f}, m(-1)/m(1) = {mdl.q(-1.0, Kc)/mdl.q(1.0, Kc):.3f}") +ax[0].set_title(r"$\beta\, u'(C_1(\epsilon))/u'(C_0)$") +ax[1].set_title(r"Arrow price density $q(\epsilon;K)$") +for a_ in ax: + a_.set_xlabel(r'$\epsilon$') + a_.legend() +plt.tight_layout() +plt.show() +``` + +**Wealth distribution.** + +From the formula for $\eta$ in `opt_c`, $\eta$ is the ratio of agent 1's wealth $w_0^1 + \theta_0^1 V + \int w_1^1(\epsilon) q(\epsilon) d\epsilon$ to aggregate wealth. + +Consequently, any rearrangement of endowments that leaves each agent's present value unchanged leaves $\eta = 0.51441$ unchanged, and with it the implied Pareto weight $\phi_1 = \eta^\psi / (\eta^\psi + (1-\eta)^\psi) = 0.54314$. + +Handing agent 1 all shares without an offsetting transfer raises agent 1's wealth by $0.5 V$ and raises $\eta$ to $0.53155$ and $\phi_1$ to $0.59365$. + +In all four cases $K = 0.14235$: as the planning problem showed, $K$ does not depend on the distribution of wealth. + +**Portfolio indeterminacy.** + +Each of the four portfolios costs $0.091850$ at time $0$ and leaves agent 1 with the same $c_0^1 = 0.955600$. + +Holding more equity simply means buying fewer Arrow securities, because equity is itself a bundle of Arrow securities priced at {eq}`tildeV2`. + +**Pricing kernel.** + +With $\chi_2 = 0.9$, aggregate time $1$ consumption rises with $\epsilon$, so the discount factor falls steeply: it is $1.309$ at $\epsilon = -1$ and $0.038$ at $\epsilon = 1$, a ratio of about $35$. + +Claims that pay in bad states are therefore expensive, which is why capital carries a large risk premium in {ref}`bcgc_ex2`. + +With $\chi_2 = -0.9$, agent 2's endowment is high when $\epsilon$ is low and output is high when $\epsilon$ is high, so aggregate consumption is high at both ends: the discount factor is hump-shaped, $0.140$ at $\epsilon = -1$, $0.323$ at $\epsilon = 0$ and $0.152$ at $\epsilon = 1$. + +The Arrow price density $q(\epsilon;K)$ multiplies the discount factor by the normal density $g(\epsilon)$, so it is single-peaked in both cases; its peak is near $\epsilon = -0.28$ when $\chi_2 = 0.9$ but near $\epsilon = -0.01$ (close to $\mu$) when $\chi_2 = -0.9$. + +```{solution-end} +``` diff --git a/lectures/BCG_incomplete_mkts.md b/lectures/BCG_incomplete_mkts.md index c6a282bb..36a29c01 100644 --- a/lectures/BCG_incomplete_mkts.md +++ b/lectures/BCG_incomplete_mkts.md @@ -33,51 +33,44 @@ In addition to what's in Anaconda, this lecture will need the following librarie ## Introduction -This is an extension of an earlier lecture {doc}`Irrelevance of Capital Structure with Complete Markets ` about a **complete markets** -model. +This is an extension of an earlier lecture {doc}`Irrelevance of Capital Structures with Complete Markets ` about a **complete markets** model. -In contrast to that lecture, this one describes an instance of a model authored by Bisin, Clementi, and Gottardi {cite}`BCG_2018` -in which financial markets are **incomplete**. +In contrast to that lecture, this one describes an instance of a model authored by Bisin, Clementi, and Gottardi {cite}`BCG_2018` in which financial markets are **incomplete**. -Instead of being able to trade equities and a full set of one-period -Arrow securities as they can in {doc}`Irrelevance of Capital Structure with Complete Markets `, here consumers and firms trade only equity and a bond. +Instead of being able to trade equities and a full set of one-period Arrow securities as they can in {doc}`Irrelevance of Capital Structures with Complete Markets `, here consumers and firms trade only equity and a bond. -It is useful to watch how outcomes differ in the two settings. +It is useful to watch how outcomes differ in the two settings. -In the complete markets economy in {doc}`Irrelevance of Capital Structure with Complete Markets ` +In the complete markets economy in {doc}`Irrelevance of Capital Structures with Complete Markets ` -- there is a unique stochastic discount factor that prices all assets +- there is a unique stochastic discount factor that prices all assets - consumers’ portfolio choices are indeterminate - firms' financial structures are indeterminate, so the model embodies an instance of a Modigliani-Miller irrelevance theorem {cite}`Modigliani_Miller_1958` -- the aggregate of all firms' financial structures are indeterminate, a consequence of there being redundant assets +- the aggregate of all firms' financial structures is indeterminate, a consequence of there being redundant assets In the incomplete markets economy studied here -- there is a not a unique equilibrium stochastic discount factor +- there is not a unique equilibrium stochastic discount factor - different stochastic discount factors price different assets - consumers’ portfolio choices are determinate - while **individual** firms' financial structures are indeterminate, thus conforming to part of a Modigliani-Miller theorem, - {cite}`Modigliani_Miller_1958`, the **aggregate** of all firms' financial structures **is** determinate. + {cite}`Modigliani_Miller_1958`, the **aggregate** of all firms' financial structures **is** determinate. -A `Big K, little k` analysis played an important role in the previous lecture {doc}`Irrelevance of Capital Structure with Complete Markets `. +A `Big K, little k` analysis played an important role in the previous lecture {doc}`Irrelevance of Capital Structures with Complete Markets `. -A more subtle version of a `Big K, little k` features in the BCG incomplete markets environment here. +A more subtle version of a `Big K, little k` features in the BCG incomplete markets environment here. -We use it to convey the heart of what BCG call a **rational conjectures** equilibrium in which conjectures are about -equilibrium pricing functions in regions of the state space that an average consumer or firm does not visit in equilibrium. +We use it to convey the heart of what BCG call a **rational conjectures** equilibrium in which conjectures are about equilibrium pricing functions in regions of the state space that an average consumer or firm does not visit in equilibrium. -Note that the absence of complete markets means that now we cannot compute competitive equilibrium prices and allocations by first solving the simple planning problem that we did in {doc}`Irrelevance of Capital Structure with Complete Markets `. +Note that the absence of complete markets means that now we cannot compute competitive equilibrium prices and allocations by first solving the simple planning problem that we did in {doc}`Irrelevance of Capital Structures with Complete Markets `. Instead, we compute an equilibrium by solving a system of simultaneous inequalities. -(Here we do not address the interesting question of whether there is a *different* planning problem that we could use to compute a -competitive equlibrium allocation.) +(Here we do not address the interesting question of whether there is a *different* planning problem that we could use to compute a competitive equilibrium allocation.) ### Setup -We adopt specifications of preferences and technologies used by Bisin, -Clemente, and Gottardi (2018) {cite}`BCG_2018` and in our earlier lecture on a complete markets -version of their model. +We adopt specifications of preferences and technologies used by Bisin, Clementi, and Gottardi {cite}`BCG_2018` and in our earlier lecture on a complete markets version of their model. The economy lasts for two periods, $t=0, 1$. @@ -93,17 +86,13 @@ A scalar random variable $\epsilon$ affects both ### Ownership -A consumer of type $i$ is endowed with $w_0^i$ units of the -time $0$ good and $w_1^i(\epsilon)$ of the time $1$ -good when the random variable takes value $\epsilon$. +A consumer of type $i$ is endowed with $w_0^i$ units of the time $0$ good and $w_1^i(\epsilon)$ of the time $1$ good when the random variable takes value $\epsilon$. -At the start of period $0$, a consumer of type $i$ also owns -$\theta^i_0$ shares of a representative firm. +At the start of period $0$, a consumer of type $i$ also owns $\theta^i_0$ shares of a representative firm. ### Measures of agents and firms -As in the companion lecture {doc}`Irrelevance of Capital Structure with Complete Markets ` that studies a complete markets version of -the model, we follow BCG in assuming that there are unit measures of +As in the companion lecture {doc}`Irrelevance of Capital Structures with Complete Markets ` that studies a complete markets version of the model, we follow BCG in assuming that there are unit measures of - consumers of type $i=1$ - consumers of type $i=2$ @@ -112,8 +101,7 @@ the model, we follow BCG in assuming that there are unit measures of $A k^\alpha e^\epsilon$ units of the time $1$ good in random state $\epsilon$ -Thus, let $\omega \in [0,1]$ index a particular consumer of type -$i$. +Thus, let $\omega \in [0,1]$ index a particular consumer of type $i$. Then define Big $C^i$ as @@ -130,9 +118,7 @@ C^i_1(\epsilon) & = \int_0^1 c^i_1(\epsilon;\omega) d \, \omega \end{aligned} $$ -In the same spirit, let $\zeta \in [0,1]$ index a particular firm -and let firm $\zeta$ purchase $k(\zeta)$ units of capital -and issue $b(\zeta)$ bonds. +In the same spirit, let $\zeta \in [0,1]$ index a particular firm and let firm $\zeta$ purchase $k(\zeta)$ units of capital and issue $b(\zeta)$ bonds. Then define Big $K$ and Big $B$ as @@ -140,22 +126,17 @@ $$ K = \int_0^1 k(\zeta) d \, \zeta, \quad B = \int_0^1 b(\zeta) d \, \zeta $$ -The assumption that there are equal measures of our three types of -agents justifies our assumption that each individual agent is a -powerless **price taker**: +The assumption that there are equal measures of our three types of agents justifies our assumption that each individual agent is a powerless **price taker**: - an individual consumer chooses its own (infinitesimal) part $c^i(\omega)$ of $C^i$ taking prices as given -- an individual firm chooses its own (infinitesmimal) part +- an individual firm chooses its own (infinitesimal) part $k(\zeta)$ of $K$ and $b(\zeta)$ of $B$ taking pricing functions as given - However, equilibrium prices depend on the `Big K, Big B, Big C` objects $K$, $B$, and $C$ -The assumption about measures of agents is a powerful device for making -a host of competitive agents take as given the equilibrium prices that -turn out to be determined by the decisions of hosts of agents who are just like -them. +The assumption about measures of agents is a powerful device for making a host of competitive agents take as given the equilibrium prices that turn out to be determined by the decisions of hosts of agents who are just like them. We call an equilibrium **symmetric** if @@ -165,12 +146,11 @@ We call an equilibrium **symmetric** if $k(\zeta) = K$, $b(\zeta) = B$ for all $\zeta \in [0,1]$ -In this lecture, we restrict ourselves to describing symmetric -equilibria. +In this lecture, we restrict ourselves to describing symmetric equilibria. ### Endowments -Per capital economy-wide endowments in periods $0$ and $1$ are +Aggregate endowments in periods $0$ and $1$ are $$ \begin{aligned} @@ -179,9 +159,7 @@ w_1(\epsilon) & = w_1^1(\epsilon) + w_1^2(\epsilon) \textrm{ in state }\epsilon \end{aligned} $$ -### Feasibility: - -Where $\alpha \in (0,1)$ and $A >0$ +### Feasibility $$ \begin{aligned} @@ -204,16 +182,11 @@ w_1^i(\epsilon) & = e^{- \chi_i \mu - .5 \chi_i^2 \sigma^2 + \chi_i \epsilon} , \end{aligned} $$ -Sometimes instead of asuming $\epsilon \sim g(\epsilon) = {\mathcal N}(0,\sigma^2)$, -we’ll assume that $g(\cdot)$ is a probability -mass function that serves as a discrete approximation to a standardized -normal density. +In the computations below, $g$ is the density of ${\mathcal N}(\mu,\sigma^2)$ truncated to $[-3, 3]$, an interval so wide (about $\pm 7.5\sigma$ at the default parameter values) that the truncation is immaterial. -### Preferences: +### Preferences -A consumer of type $i$ orders period $0$ consumption -$c_0^i$ and state $\epsilon$-period $1$ consumption -$c^i(\epsilon)$ by +A consumer of type $i$ orders period $0$ consumption $c_0^i$ and period $1$, state $\epsilon$ consumption $c^i_1(\epsilon)$ by $$ u^i = u(c_0^i) + \beta \int u(c_1^i(\epsilon)) g (\epsilon) d \epsilon, \quad i = 1,2 @@ -230,39 +203,25 @@ $$ ### Risk-sharing motives -The two types of agents’ period $1$ endowments have different correlations with -the physical return on capital. +The two types of agents’ period $1$ endowments have different correlations with the physical return on capital. -Endowment differences give agents incentives to trade risks that in the -complete market version of the model showed up in their demands for -equity and in their demands and supplies of one-period Arrow securities. +Endowment differences give agents incentives to trade risks that in the complete market version of the model showed up in their demands for equity and in their demands and supplies of one-period Arrow securities. -In the incomplete-markets setting under study here, these differences -show up in differences in the two types of consumers’ demands for a -typical firm’s bonds and equity, the only two assets that agents can now -trade. +In the incomplete-markets setting under study here, these differences show up in differences in the two types of consumers’ demands for a typical firm’s bonds and equity, the only two assets that agents can now trade. ## Asset markets -Markets are incomplete: *ex cathedra* we the model builders declare that only equities and bonds issued by representative -firms can be traded. +Markets are incomplete: *ex cathedra* we the model builders declare that only equities and bonds issued by representative firms can be traded. -Let $\theta^i$ and $\xi^i$ be a consumer of type -$i$’s post-trade holdings of equity and bonds, respectively. +Let $\theta^i$ and $\xi^i$ be a consumer of type $i$’s post-trade holdings of equity and bonds, respectively. -A firm issues bonds promising to pay $b$ units of consumption at -time $t=1$ and purchases $k$ units of physical capital at -time $t=0$. +A firm issues bonds promising to pay $b$ units of consumption at time $t=1$ and purchases $k$ units of physical capital at time $t=0$. -When $e^\epsilon A k^\alpha < b$ at time $1$, the firm defaults and its output is -divided equally among bondholders. +When $e^\epsilon A k^\alpha < b$ at time $1$, the firm defaults and its output is divided equally among bondholders. -Evidently, when the productivity shock -$\epsilon < \epsilon^* = \log \left(\frac{b}{ Ak^\alpha}\right)$, -the firm defaults on its debt +Evidently, when the productivity shock $\epsilon < \epsilon^* = \log \left(\frac{b}{ Ak^\alpha}\right)$, the firm defaults on its debt. -Payoffs to equity and debt at date 1 as functions of the productivity -shock $\epsilon$ are thus +Payoffs to equity and debt at date 1 as functions of the productivity shock $\epsilon$ are thus ```{math} :label: payofffns @@ -273,41 +232,29 @@ d^b(k,b;\epsilon) &= \min \left\{ \frac{e^\epsilon A k^\alpha}{b}, 1 \right\} \end{aligned} ``` -A firm faces a bond price function $p(k,b)$ when it issues -$b$ bonds and purchases $k$ units of physical capital. +A firm faces a bond price function $p(k,b)$ when it issues $b$ bonds and purchases $k$ units of physical capital. + +A firm’s equity is worth $q(k,b)$ when it issues $b$ bonds and purchases $k$ units of physical capital. -A firm’s equity is worth $q(k,b)$ when it issues $b$ bonds -and purchases $k$ units of physical capital. +In {doc}`the companion lecture `, $q(\epsilon)$ instead denoted the price of an Arrow security paying in state $\epsilon$; no such securities are traded in the present economy, where markets are incomplete. -A firm regards an equity-pricing function $q(k,b)$ and a bond -pricing function $p(k,b)$ as exogenous in the sense that they are -not affected by its choices of $k$ and $b$. +A firm regards an equity-pricing function $q(k,b)$ and a bond pricing function $p(k,b)$ as exogenous in the sense that they are not affected by its choices of $k$ and $b$. -Consumers face equilibrium prices $\check q$ and $\check p$ -for bonds and equities, where $\check q$ and $\check p$ are -both scalars. +Consumers face equilibrium prices $\check q$ and $\check p$ for equities and bonds, respectively, where $\check q$ and $\check p$ are both scalars. Consumers are price takers and only need to know the scalars $\check q, \check p$. -Firms are *price function* takers and must know the functions $q(k,b), p(k,b)$ in order -completely to pose their optimum problems. +Firms are *price function* takers and must know the functions $q(k,b), p(k,b)$ in order completely to pose their optimum problems. ### Consumers -Each consumer of type $i$ is endowed with $w_0^i$ of the -time $0$ consumption good, $w_1^i(\epsilon)$ of the time -$1$, state $\epsilon$ consumption good and also owns a fraction -$\theta^i_0 \in (0,1)$ of the initial value of a representative -firm, where $\theta^1_0 + \theta^2_0 = 1$. +Each consumer of type $i$ is endowed with $w_0^i$ of the time $0$ consumption good, $w_1^i(\epsilon)$ of the time $1$, state $\epsilon$ consumption good and also owns a fraction $\theta^i_0 \in (0,1)$ of the initial value of a representative firm, where $\theta^1_0 + \theta^2_0 = 1$. -The initial value of a representative firm is $V$ (an object to be -determined in a rational expectations equilibrium). +The initial value of a representative firm is $V$ (an object to be determined in a rational expectations equilibrium). -Consumer $i$ buys $\theta^i$ shares of equity and buys bonds -worth $\check p \xi^i$ where $\check p$ is the bond price. +Consumer $i$ buys $\theta^i$ shares of equity and buys bonds worth $\check p \xi^i$ where $\check p$ is the bond price. -Being a price-taker, a consumer takes $V$, $\check q$, $\check p$, and $K, B$ -as given. +Being a price-taker, a consumer takes $V$, $\check q$, $\check p$, and $K, B$ as given. Consumers know that equilibrium payoff functions for bonds and equities take the form @@ -322,7 +269,7 @@ Consumer $i$’s optimization problem is $$ \begin{aligned} -\max_{c^i_0,\theta^i,\xi^i,c^i_1(\epsilon)} & u(c^i_0) + \beta \int u(c^i(\epsilon)) g(\epsilon) \ d\epsilon \\ +\max_{c^i_0,\theta^i,\xi^i,c^i_1(\epsilon)} & u(c^i_0) + \beta \int u(c^i_1(\epsilon)) g(\epsilon) \ d\epsilon \\ \text{subject to } \quad & c^i_0 = w^i_0 + \theta^i_0V - \check q\theta^i - \check p \xi^i, \\ & c^i_1(\epsilon) = w^i_1(\epsilon) + \theta^i d^e(K,B;\epsilon) + \xi^i d^b(K,B;\epsilon) \ \forall \ \epsilon, \\ @@ -330,45 +277,41 @@ $$ \end{aligned} $$ -The last two inequalities impose that the consumer cannot short sell either -equity or bonds. +The last two inequalities impose that the consumer cannot short sell either equity or bonds. -In a rational expectations equilibrium, $\check q = q(K,B)$ and $\check p = p(K,B)$ +In a rational expectations equilibrium, $\check q = q(K,B)$ and $\check p = p(K,B)$. We form consumer $i$’s Lagrangian: $$ \begin{aligned} -L^i := & u(c^i_0) + \beta \int u(c^i(\epsilon)) g(\epsilon) \ d\epsilon \\ - & +\lambda^i_0 [w^i_0 + \theta_0V - \check q\theta^i - \check p \xi^i - c^i_0] \\ +L^i := & u(c^i_0) + \beta \int u(c^i_1(\epsilon)) g(\epsilon) \ d\epsilon \\ + & +\lambda^i_0 [w^i_0 + \theta^i_0 V - \check q\theta^i - \check p \xi^i - c^i_0] \\ & + \beta \int \lambda^i_1(\epsilon) \left[ w^i_1(\epsilon) + \theta^i d^e(K,B;\epsilon) + \xi^i d^b(K,B;\epsilon) - c^i_1(\epsilon) \right] g(\epsilon) \ d\epsilon \end{aligned} $$ -Consumer $i$’s first-order necessary conditions for an optimum -include: +Consumer $i$’s first-order necessary conditions for an optimum include: $$ \begin{aligned} c^i_0:& \quad u^\prime(c^i_0) = \lambda^i_0 \\ c^i_1(\epsilon):& \quad u^\prime(c^i_1(\epsilon)) = \lambda^i_1(\epsilon) \\ \theta^i:& \quad \beta \int \lambda^i_1(\epsilon) d^e(K,B;\epsilon) g(\epsilon) \ d\epsilon \leq \lambda^i_0 \check q \quad (= \ \ \text{if} \ \ \theta^i>0) \\ -\xi^i:& \quad \beta \int \lambda^i_1(\epsilon) d^b(K,B;\epsilon) g(\epsilon) \ d\epsilon \leq \lambda^i_0 \check p \quad (= \ \ \text{if} \ \ b^i>0) \\ +\xi^i:& \quad \beta \int \lambda^i_1(\epsilon) d^b(K,B;\epsilon) g(\epsilon) \ d\epsilon \leq \lambda^i_0 \check p \quad (= \ \ \text{if} \ \ \xi^i>0) \\ \end{aligned} $$ -We can combine and rearrange consumer $i$’s first-order -conditions to become: +We can combine and rearrange consumer $i$’s first-order conditions to become: $$ \begin{aligned} \check q \geq \beta \int \frac{u^\prime(c^i_1(\epsilon))}{u^\prime(c^i_0)} d^e(K,B;\epsilon) g(\epsilon) \ d\epsilon \quad (= \ \ \text{if} \ \ \theta^i>0) \\ -\check p \geq \beta \int \frac{u^\prime(c^i_1(\epsilon))}{u^\prime(c^i_0)} d^b(K,B;\epsilon) g(\epsilon) \ d\epsilon \quad (= \ \ \text{if} \ \ b^i>0)\\ +\check p \geq \beta \int \frac{u^\prime(c^i_1(\epsilon))}{u^\prime(c^i_0)} d^b(K,B;\epsilon) g(\epsilon) \ d\epsilon \quad (= \ \ \text{if} \ \ \xi^i>0)\\ \end{aligned} $$ -These inequalities imply that in a symmetric rational expectations equilibrium consumption allocations and -prices satisfy +These inequalities imply that in a symmetric rational expectations equilibrium consumption allocations and prices satisfy $$ \begin{aligned} @@ -379,12 +322,9 @@ $$ ### Pricing functions -When individual firms solve their optimization problems, they take big -$C^i$’s as fixed objects that they don’t influence. +When individual firms solve their optimization problems, they take big $C^i$’s as fixed objects that they don’t influence. -A representative firm faces a price function $q(k,b)$ for its -equity and a price function $p(k, b)$ per unit of bonds that -satisfy +A representative firm faces a price function $q(k,b)$ for its equity and a price function $p(k, b)$ per unit of bonds that satisfy $$ \begin{aligned} @@ -395,50 +335,39 @@ $$ where the payoff functions are described by equations {eq}`payofffns`. -Notice the appearance of big $C^i$’s on the right sides of these -two equations that define equilibrium pricing functions. +Notice the appearance of big $C^i$’s on the right sides of these two equations that define equilibrium pricing functions. -The two price functions describe outcomes not only for equilibrium choices -$K, B$ of capital $k$ and debt $b$, but also for any -**out-of-equilibrium** pairs $(k, b) \neq (K, B)$. +The two price functions describe outcomes not only for equilibrium choices $K, B$ of capital $k$ and debt $b$, but also for any **out-of-equilibrium** pairs $(k, b) \neq (K, B)$. The firm is assumed to know both price functions. This means that the firm understands that its choice of $k,b$ influences how markets price its equity and debt. -This package of assumptions is sometimes called **rational conjectures** (about price functions). +This package of assumptions is sometimes called **rational conjectures** (about price functions). -BCG give credit to Makowski for emphasizing and clarifying how rational conjectures are components of rational expectations equilibria. +BCG give credit to Makowski {cite}`Makowski_1983` for emphasizing and clarifying how rational conjectures are components of rational expectations equilibria. ### Firms -The firm chooses capital $k$ and debt $b$ to maximize its -market value: +The firm chooses capital $k$ and debt $b$ to maximize its market value: $$ V \equiv \max_{k,b} -k + q(k,b) + p(k,b) b $$ -Attributing value maximization to the firm is a good idea because in equilibrium consumers of both types -*want* a firm to maximize its value. +Attributing value maximization to the firm is a good idea because in equilibrium consumers of both types *want* a firm to maximize its value. In the special quantitative examples studied here -- consumers of types $i=1,2$ both hold equity +- consumers of types $i=1,2$ both hold equity - only consumers of type $i=2$ hold debt; consumers of type $i=1$ hold none. -These outcomes occur because we follow BCG and set parameters so that a -type 2 consumer’s stochastic endowment of the consumption good in period -$1$ is more correlated with the firm’s output than is a type 1 -consumer’s. +These outcomes occur because we follow BCG and set parameters so that a type 2 consumer’s stochastic endowment of the consumption good in period $1$ is more correlated with the firm’s output than is a type 1 consumer’s. -This gives consumers of type $2$ a motive to hedge their second period -endowment risk by holding bonds (they also choose to -hold some equity). +This gives consumers of type $2$ a motive to hedge their second period endowment risk by holding bonds (they also choose to hold some equity). -These outcomes mean that the pricing functions end up -satisfying +These outcomes mean that the pricing functions end up satisfying $$ \begin{aligned} @@ -447,9 +376,7 @@ p(k,b) &= \beta \int \frac{u^\prime(C^2_1(\epsilon))}{u^\prime(C^2_0)} d^b(k,b;\ \end{aligned} $$ -Recall that -$\epsilon^*(k,b) \equiv \log\left(\frac{b}{Ak^\alpha}\right)$ is a -firm’s default threshold. +Recall that $\epsilon^*(k,b) \equiv \log\left(\frac{b}{Ak^\alpha}\right)$ is a firm’s default threshold. We can rewrite the pricing functions as: @@ -468,18 +395,16 @@ $$ V \equiv \max_{k,b} \left\{ -k + q(k,b) + p(k, b) b \right\} $$ -The firm’s first-order necessary conditions with respect to $k$ -and $b$, respectively, are +The firm’s first-order necessary conditions with respect to $k$ and $b$, respectively, are $$ \begin{aligned} -k: \quad & -1 + \frac{\partial q(k,b)}{\partial k} + b \frac{\partial p(q,b)}{\partial k} = 0 \cr +k: \quad & -1 + \frac{\partial q(k,b)}{\partial k} + b \frac{\partial p(k,b)}{\partial k} = 0 \cr b: \quad & \frac{\partial q(k,b)}{\partial b} + p(k,b) + b \frac{\partial p(k,b)}{\partial b} = 0 \end{aligned} $$ -We use the Leibniz integral rule several times to arrive at -the following derivatives: +We use the Leibniz integral rule several times to arrive at the following derivatives: $$ \frac{\partial q(k,b)}{\partial k} = \beta \alpha A k^{\alpha-1} \int_{\epsilon^*}^\infty \frac{u'(C_1^i(\epsilon))}{u'(C_0^i)} @@ -491,22 +416,30 @@ $$ $$ $$ -\frac{\partial p(k,b)}{\partial k} = \beta \alpha \frac{A k^{\alpha -1}}{b} \int_{-\infty}^{\epsilon^*} \frac{u'(C_1^2(\epsilon))}{u'(C_0^2)} g(\epsilon) d \epsilon +\frac{\partial p(k,b)}{\partial k} = \beta \alpha \frac{A k^{\alpha -1}}{b} \int_{-\infty}^{\epsilon^*} \frac{u'(C_1^2(\epsilon))}{u'(C_0^2)} e^\epsilon g(\epsilon) d \epsilon $$ $$ \frac{\partial p(k,b)}{\partial b} = - \beta \frac{A k^\alpha}{b^2} \int_{-\infty}^{\epsilon^*} \frac{u'(C_1^2(\epsilon))}{u'(C_0^2)} e^\epsilon g(\epsilon) d \epsilon $$ -**Special case:** We confine ourselves to a special case in which both types of -consumer hold positive equities so that -$\frac{\partial q(k,b)}{\partial k}$ and -$\frac{\partial q(k,b)}{\partial b}$ are related to rates of -intertemporal substitution for both agents. +Each expression labeled $i=1,2$ is the derivative of type $i$'s own valuation of equity, + +$$ +Q^i(k,b) = \beta \int_{\epsilon^*}^\infty \frac{u^\prime(C^i_1(\epsilon))}{u^\prime(C^i_0)} \left( e^\epsilon Ak^\alpha - b \right) g(\epsilon) \ d\epsilon , +$$ + +and the two expressions agree only where both types value equity equally. + +That is why, in the special case described next, {eq}`Eqn2` below is not implied by the two formulas but is a separate requirement: at the equilibrium $(K,B)$, the valuations $Q^1$ and $Q^2$ must change at the same rate as a firm varies $b$, so that type $1$ remains willing to hold equity. + +Exercise {ref}`bcgi_ex3` verifies this numerically. + +**Special case:** We confine ourselves to a special case in which both types of consumer hold positive equities so that $\frac{\partial q(k,b)}{\partial k}$ and $\frac{\partial q(k,b)}{\partial b}$ are related to rates of intertemporal substitution for both agents. + +The code below computes only equilibria of this special case, and it prints a warning when a solution would require one type to hold no equity. -Substituting these partial derivatives into the above first-order -conditions for $k$ and $b$, respectively, we obtain the -following versions of those first order conditions: +Substituting these partial derivatives into the above first-order conditions for $k$ and $b$, respectively, we obtain the following versions of those first order conditions: ```{math} :label: Eqn1 @@ -521,28 +454,20 @@ b: \quad \int_{\epsilon^*}^\infty \left( \frac{u^\prime(C^1_1(\epsilon))}{u^\prime(C^1_0)} \right) g(\epsilon) \ d\epsilon = \int_{\epsilon^*}^\infty \left( \frac{u^\prime(C^2_1(\epsilon))}{u^\prime(C^2_0)} \right) g(\epsilon) \ d\epsilon ``` -where again recall that -$\epsilon^*(k,b) \equiv \log\left(\frac{b}{Ak^\alpha}\right)$. +where again recall that $\epsilon^*(k,b) \equiv \log\left(\frac{b}{Ak^\alpha}\right)$. -Taking $C_0^i, C_1^i(\epsilon)$ as given, these are two equations -that we want to solve for the firm’s optimal decisions $k, b$. +Taking $C_0^i, C_1^i(\epsilon)$ as given, these are two equations that we want to solve for the firm’s optimal decisions $k, b$. ## Equilibrium verification -On page 5 of BCG (2018), the authors say +On page 5 of {cite:t}`BCG_2018`, the authors say -*If the price conjectures corresponding to the plan chosen by firms in -equilibrium are correct, that is equal to the market prices* $\check q$ *and* $\check p$, *it is immediate to verify that -the rationality of the conjecture coincides with the agents’ Euler -equations.* +*If the price conjectures corresponding to the plan chosen by firms in equilibrium are correct, that is equal to the market prices* $\check q$ *and* $\check p$, *it is immediate to verify that the rationality of the conjecture coincides with the agents’ Euler equations.* -Here BCG are describing how they go about verifying that when they set -little $k$, little $b$ from the firm’s first-order -conditions equal to the big $K$, big $B$ at the big -$C$’s that appear in the pricing functions, then +Here BCG are describing how they go about verifying that when they set little $k$, little $b$ from the firm’s first-order conditions equal to the big $K$, big $B$ at the big $C$’s that appear in the pricing functions, then - consumers’ Euler equations are satisfied if little $c$’s are - equated to Big $C$’s + equated to Big $C$’s - firms’ first-order necessary conditions for $k, b$ are satisfied. - $\check q = q(K,B)$ and @@ -550,8 +475,7 @@ $C$’s that appear in the pricing functions, then ## Pseudo code -Before displaying our Python code for computing a BCG incomplete markets equilibrium, -we’ll sketch some pseudo code that describes its logical flow. +Before displaying our Python code for computing a BCG incomplete markets equilibrium, we’ll sketch some pseudo code that describes its logical flow. Here goes: @@ -565,17 +489,17 @@ Here goes: $\epsilon^* \equiv \log\left(\frac{b}{Ak^\alpha}\right)$. 1. (In this step we abuse notation by freezing $V, k, b$ and in effect temporarily treating them as Big $K,B$ values. Thus, in - this step 6 little $k, b$ are frozen at guessed at value of $K, B$.) + this step 6 little $k, b$ are frozen at guessed values of $K, B$.) Fixing the values of $V$, $b$ and $k$, compute optimal choices of consumption $c^i$ with consumers’ FOCs. Assume that only agent 2 holds debt: $\xi^2 = b$ and that both agents hold equity: $0 <\theta^i < 1$ for $i=1,2$. -1. Set high and low bounds for equity holdings for agent 1 as $\theta^1_h$ and $\theta^1_l$. Guess +1. Set high and low bounds for equity holdings for agent 1 as $\theta^1_h$ and $\theta^1_l$. Guess $\theta^1 = \frac{1}{2}(\theta^1_h + \theta^1_l)$, and $\theta^2 = 1 - \theta^1$. While $|\theta^1_h - \theta^1_l|$ is large: - * Compute agent 1’s valuation of the equity claim with a - fixed-point iteration: + * Compute agent 1’s valuation of the equity claim by + bisection: $q_1 = \beta \int \frac{u^\prime(c^1_1(\epsilon))}{u^\prime(c^1_0)} d^e(k,b;\epsilon) g(\epsilon) \ d\epsilon$ @@ -586,35 +510,35 @@ Here goes: and $c^1_0 = w^1_0 + \theta^1_0V - q_1\theta^1$ - * Compute agent 2’s valuation of the bond claim with a - fixed-point iteration: + * Compute agent 2’s valuation of the bond claim by + bisection: $p = \beta \int \frac{u^\prime(c^2_1(\epsilon))}{u^\prime(c^2_0)} d^b(k,b;\epsilon) g(\epsilon) \ d\epsilon$ where - $c^2_1(\epsilon) = w^2_1(\epsilon) + \theta^2 d^e(k,b;\epsilon) + b$ + $c^2_1(\epsilon) = w^2_1(\epsilon) + \theta^2 d^e(k,b;\epsilon) + b\, d^b(k,b;\epsilon)$ and $c^2_0 = w^2_0 + \theta^2_0 V - q_1 \theta^2 - pb$ - * Compute agent 2’s valuation of the equity claim with a - fixed-point iteration: + * Compute agent 2’s valuation of the equity claim by + bisection: $q_2 = \beta \int \frac{u^\prime(c^2_1(\epsilon))}{u^\prime(c^2_0)} d^e(k,b;\epsilon) g(\epsilon) \ d\epsilon$ where - $c^2_1(\epsilon) = w^2_1(\epsilon) + \theta^2 d^e(k,b;\epsilon) + b$ + $c^2_1(\epsilon) = w^2_1(\epsilon) + \theta^2 d^e(k,b;\epsilon) + b\, d^b(k,b;\epsilon)$ and $c^2_0 = w^2_0 + \theta^2_0 V - q_2 \theta^2 - pb$ - * If $q_1 > q_2$, Set $\theta_l = \theta^1$; - otherwise, set $\theta_h = \theta^1$. - * Repeat steps 6Aa through 6Ad until + * If $q_1 > q_2$, Set $\theta^1_l = \theta^1$; + otherwise, set $\theta^1_h = \theta^1$. + * Repeat the four sub-steps above until $|\theta^1_h - \theta^1_l|$ is small. -1. Set bond price as $p$ and equity price as $q = \max(q_1,q_2)$. +1. Set bond price as $p$ and equity price as $q = \max(q_1,q_2)$. 1. Compute optimal choices of consumption: $$ @@ -622,51 +546,49 @@ Here goes: c^1_0 &= w^1_0 + \theta^1_0V - q\theta^1 \\ c^2_0 &= w^2_0 + \theta^2_0V - q\theta^2 - pb \\ c^1_1(\epsilon) &= w^1_1(\epsilon) + \theta^1 d^e(k,b;\epsilon) \\ - c^2_1(\epsilon) &= w^2_1(\epsilon) + \theta^2 d^e(k,b;\epsilon) + b + c^2_1(\epsilon) &= w^2_1(\epsilon) + \theta^2 d^e(k,b;\epsilon) + b\, d^b(k,b;\epsilon) \end{aligned} $$ 1. (Here we confess to abusing notation again, but now in a different - way. In step 7, we interpret frozen $c^i$s as Big + way. In steps 10 through 12, we interpret frozen $c^i$s as Big $C^i$. We do this to solve the firm’s problem.) Fixing the values of $c^i_0$ and $c^i_1(\epsilon)$, compute optimal choices of capital $k$ and debt level $b$ using the firm’s first order necessary conditions. 1. Compute deviations from the firm’s FONC for capital $k$ as: - $kfoc = \beta \alpha A k^{\alpha - 1} \left( \int \frac{u^\prime(c^2_1(\epsilon))}{u^\prime(c^2_0)} e^\epsilon g(\epsilon) \ d\epsilon \right) - 1$ + $kfoc = \beta \alpha A k^{\alpha - 1} \left( \int \frac{u^\prime(c^2_1(\epsilon))}{u^\prime(c^2_0)} e^\epsilon g(\epsilon) \ d\epsilon \right) - 1$ - If $kfoc > 0$, Set $k_l = k$; otherwise, set $k_h = k$. - - Repeat steps 4 through 7A until $|k_h-k_l|$ is small. -1. Compute deviations from the firm’s FONC for debt level $b$ as: + - Repeat steps 4 through 11 until $|k_h-k_l|$ is small. +1. Compute deviations from the firm’s FONC for debt level $b$ as: - $bfoc = \beta \left[ \int_{\epsilon^*}^\infty \left( \frac{u^\prime(c^1_1(\epsilon))}{u^\prime(c^1_0)} \right) g(\epsilon) \ d\epsilon - \int_{\epsilon^*}^\infty \left( \frac{u^\prime(c^2_1(\epsilon))}{u^\prime(c^2_0)} \right) g(\epsilon) \ d\epsilon \right]$ + $bfoc = \beta \left[ \int_{\epsilon^*}^\infty \left( \frac{u^\prime(c^1_1(\epsilon))}{u^\prime(c^1_0)} \right) g(\epsilon) \ d\epsilon - \int_{\epsilon^*}^\infty \left( \frac{u^\prime(c^2_1(\epsilon))}{u^\prime(c^2_0)} \right) g(\epsilon) \ d\epsilon \right]$ - If $bfoc > 0$, Set $b_h = b$; otherwise, set $b_l = b$. - - Repeat steps 3 through 7B until $|b_h-b_l|$ is small. -1. Given prices $q$ and $p$ from step 6, and the firm - choices of $k$ and $b$ from step 7, compute the synthetic + - Repeat steps 3 through 12 until $|b_h-b_l|$ is small. +1. Given prices $q$ and $p$ from step 8, and the firm + choices of $k$ and $b$ from steps 11 and 12, compute the synthetic firm value: $V_x = -k + q + pb$ - If $V_x > V$, then set $V_l = V$; otherwise, set $V_h = V$. - - Repeat steps 1 through 8 until $|V_x - V|$ is small. -1. Ultimately, the algorithm returns equilibrium capital + - Repeat steps 2 through 13 until $|V_x - V|$ is small. +1. Ultimately, the algorithm returns equilibrium capital $k^*$, debt $b^*$ and firm value $V^*$, as well as the following equilibrium values: - Equity holdings $\theta^{1,*} = \theta^1(k^*,b^*)$ - Prices $q^*=q(k^*,b^*), \ p^*=p(k^*,b^*)$ - Consumption plans - $C^{1,*}_0 = c^1_0(k^*,b^*),\ C^{2,*}_0 = c^2_0(k^*,b^*), \ C^{1,*}_1(\epsilon) = c^1_1(k^*,b^*;\epsilon),\ C^{1,*}_1(\epsilon) = c^2_1(k^*,b^*;\epsilon)$. + $C^{1,*}_0 = c^1_0(k^*,b^*),\ C^{2,*}_0 = c^2_0(k^*,b^*), \ C^{1,*}_1(\epsilon) = c^1_1(k^*,b^*;\epsilon),\ C^{2,*}_1(\epsilon) = c^2_1(k^*,b^*;\epsilon)$. ## Code -We create a Python class `BCG_incomplete_markets` to compute the -equilibrium allocations of the incomplete market BCG model, given a set -of parameter values. +We create a Python class `BCG_incomplete_markets` to compute the equilibrium allocations of the incomplete market BCG model, given a set of parameter values. -The class includes the following methods, i.e., functions: +The class includes the following methods, i.e., functions: - `solve_eq`: solves the BCG model and returns the equilibrium values of capital $k$, debt $b$ and firm value $V$, as @@ -675,7 +597,9 @@ The class includes the following methods, i.e., functions: - prices $q^*, p^*$ - consumption plans $C^{1,*}_0, C^{2,*}_0, C^{1,*}_1(\epsilon), C^{2,*}_1(\epsilon)$. -- `eq_valuation`: inputs equilibrium consumpion plans $C^*$ and +- `valuations_by_agent`: given consumption plans and a pair $(k,b)$, returns each agent's + valuations $Q^1, Q^2$ of equity and $P^1, P^2$ of bonds. +- `eq_valuation`: inputs equilibrium consumption plans $C^*$ and outputs the following valuations for each pair of $(k,b)$ in the grid: - the firm $V(k,b)$ @@ -684,20 +608,23 @@ The class includes the following methods, i.e., functions: Parameters include: -- $\chi_1$, $\chi_2$: correlation parameter for agent 1 +- $\chi_1$, $\chi_2$: correlation parameter for agent 1 and 2. Default values are respectively 0 and 0.9. -- $w^1_0$, $w^2_0$: initial endowments. Default values +- $w^1_0$, $w^2_0$: initial endowments. Default values are respectively 0.9 and 1.1. -- $\theta^1_0$, $\theta^2_0$: initial holding of the +- $\theta^1_0$, $\theta^2_0$: initial holding of the firm. Default values are 0.5. -- $\psi$: risk parameter. Default value is 3. +- $\gamma$ (`𝜓1`, `𝜓2` in the code): coefficients of relative risk aversion of agents 1 and 2. + Default values are 3. - $\alpha$: Production function parameter. Default value is 0.6. - $A$: Productivity of the firm. Default value is 2.5. - $\mu$, $\sigma$: Mean and standard deviation of the shock distribution. Default values are respectively -0.025 and 0.4 - $\beta$: Discount factor. Default value is 0.96. -- bound: Bound for truncated normal distribution. Default value is 3. +- bound: Bound, in units of $\epsilon$, for truncated normal distribution. Default value is 3. +- `Vl`, `Vh`, `kbot`, `ktop`, `bbot`, `btop`: lower and upper bisection bounds for firm value $V$, + capital $k$, and debt $b$. Default values are respectively 0, 0.5, 0.01, 0.25, 0.1, and 0.8. These bounds must bracket the solution; if they do not, the bisections will fail to converge. ```{code-cell} ipython3 import numpy as np @@ -728,7 +655,6 @@ class BCG_incomplete_markets: Vl = 0, Vh = 0.5, kbot = 0.01, - #ktop = (𝛼*A)**(1/(1-𝛼)), ktop = 0.25, bbot = 0.1, btop = 0.8): @@ -752,14 +678,10 @@ class BCG_incomplete_markets: self.Vl = Vl self.Vh = Vh self.kbot = kbot - #self.kbot = (𝛼*A)**(1/(1-𝛼)) self.ktop = ktop self.bbot = bbot self.btop = btop - # Utility - self.u = njit(lambda c: (c**(1-𝜓)) / (1-𝜓)) - # Initial endowments self.w10 = w10 self.w20 = w20 @@ -777,11 +699,10 @@ class BCG_incomplete_markets: # Truncated normal ta, tb = (-bound - 𝜇) / 𝜎, (bound - 𝜇) / 𝜎 rv = truncnorm(ta, tb, loc=𝜇, scale=𝜎) - 𝜖_range = np.linspace(ta, tb, 1000000) + 𝜖_range = np.linspace(-bound, bound, 1000000) pdf_range = rv.pdf(𝜖_range) self.g = njit(lambda 𝜖: np.interp(𝜖, 𝜖_range, pdf_range)) - #************************************************************* # Function: Solve for equilibrium of the BCG model #************************************************************* @@ -850,11 +771,10 @@ class BCG_incomplete_markets: # Production fk = A*(k**𝛼) -# Y = lambda 𝜖: np.exp(𝜖)*fk # Compute integration threshold - epstar = np.log(b/fk) - + # Default threshold, kept inside the integration range [-bound, bound] + epstar = min(max(np.log(b/fk), -bound), bound) #************************************************************** # Compute the prices and allocations consistent with consumers' @@ -883,12 +803,9 @@ class BCG_incomplete_markets: ## First, compute the constant term that is not influenced by q ## that is, 𝛽E[u'(c^{1}_{1})d^{e}(k,B)] -# intqq1 = lambda 𝜖: (w11(𝜖) + 𝜃1*(Y(𝜖, fk) - b))**(-𝜓1)*(Y(𝜖, fk) - b)*g(𝜖) -# const_qq1 = 𝛽 * quad(intqq1,epstar,bound)[0] const_qq1 = 𝛽 * quad(intqq1,epstar,bound, args=(fk, 𝜃1, 𝜓1, b))[0] - ## Second, iterate to get the equity price q qq1l = 0 qq1h = ww10 @@ -913,9 +830,6 @@ class BCG_incomplete_markets: ## First, compute the constant term that is not influenced by p ## that is, 𝛽E[u'(c^{2}_{1})d^{b}(k,B)] -# intp1 = lambda 𝜖: (Y(𝜖, fk)/b)*(w21(𝜖) + Y(𝜖, fk))**(-𝜓2)*g(𝜖) -# intp2 = lambda 𝜖: (w21(𝜖) + 𝜃2*(Y(𝜖, fk)-b) + b)**(-𝜓2)*g(𝜖) -# const_p = 𝛽 * (quad(intp1,-bound,epstar)[0] + quad(intp2,epstar,bound)[0]) const_p = 𝛽 * (quad(intp1,-bound,epstar, args=(fk, 𝜓2, b))[0]\ + quad(intp2,epstar,bound, args=(fk, 𝜃2, 𝜓2, b))[0]) @@ -933,7 +847,6 @@ class BCG_incomplete_markets: diff = abs(pl-ph) # qq2 is the equity price consistent with agent-2 Euler Equation -# intqq2 = lambda 𝜖: (w21(𝜖) + 𝜃2*(Y(𝜖, fk)-b) + b)**(-𝜓2)*(Y(𝜖, fk) - b)*g(𝜖) const_qq2 = 𝛽 * quad(intqq2,epstar,bound, args=(fk, 𝜃2, 𝜓2, b))[0] qq2l = 0 qq2h = ww20 @@ -967,18 +880,14 @@ class BCG_incomplete_markets: c20 = ww20 - q*(1-𝜃1) - p*b c21 = lambda 𝜖: w21(𝜖) + (1-𝜃1)*max(Y(𝜖, fk)-b,0) + min(Y(𝜖, fk),b) - #************************************************* # Compute the first order conditions for the firm #************************************************* - #=========== - # Equity FOC - #=========== - # Only agent 2's IMRS is relevent -# intk1 = lambda 𝜖: (w21(𝜖) + Y(𝜖, fk))**(-𝜓2)*np.exp(𝜖)*g(𝜖) -# intk2 = lambda 𝜖: (w21(𝜖) + 𝜃2*(Y(𝜖, fk)-b) + b)**(-𝜓2)*np.exp(𝜖)*g(𝜖) -# kfoc_num = quad(intk1,-bound,epstar)[0] + quad(intk2,epstar,bound)[0] + #============ + # Capital FOC + #============ + # Only agent 2's IMRS is relevant kfoc_num = quad(intk1,-bound,epstar, args=(fk, 𝜓2))[0] + quad(intk2,epstar,bound, args=(fk, 𝜃2, 𝜓2, b))[0] kfoc_denom = (ww20- q*𝜃2 - p*b)**(-𝜓2) kfoc = 𝛽*𝛼*A*(k**(𝛼-1))*(kfoc_num/kfoc_denom) - 1 @@ -992,15 +901,9 @@ class BCG_incomplete_markets: if print_crit: print("critical value of k: {:.5f}".format(k_crit)) - #========= # Bond FOC #========= -# intB1 = lambda 𝜖: (w11(𝜖) + 𝜃1*(Y(𝜖, fk) - b))**(-𝜓1)*g(𝜖) -# intB2 = lambda 𝜖: (w21(𝜖) + 𝜃2*(Y(𝜖, fk) - b) + b)**(-𝜓2)*g(𝜖) - -# bfoc1 = quad(intB1,epstar,bound)[0] / (ww10 - q*𝜃1)**(-𝜓1) -# bfoc2 = quad(intB2,epstar,bound)[0] / (ww20 - q*𝜃2 - p*b)**(-𝜓2) bfoc1 = quad(intB1,epstar,bound, args=(fk, 𝜃1, 𝜓1, b))[0] / (ww10 - q*𝜃1)**(-𝜓1) bfoc2 = quad(intB2,epstar,bound, args=(fk, 𝜃2, 𝜓2, b))[0] / (ww20 - q*𝜃2 - p*b)**(-𝜓2) @@ -1026,17 +929,17 @@ class BCG_incomplete_markets: if print_crit: print("#====== critical value of V: {:.5f}".format(V_crit)) - print('k,b,p,q,kfoc,bfoc,epstar,V,V_crit') - formattedList = ["%.3f" % member for member in [k, - b, - p, - q, - kfoc, - bfoc, - epstar, - V, - V_crit]] - print(formattedList) + print('k,b,p,q,kfoc,bfoc,epstar,V,V_crit') + formattedList = ["%.3f" % member for member in [k, + b, + p, + q, + kfoc, + bfoc, + epstar, + V, + V_crit]] + print(formattedList) #********************************* # Equilibrium values @@ -1054,24 +957,19 @@ class BCG_incomplete_markets: c21ss = c21 𝜃1ss = 𝜃1 + # The first-order conditions imposed above assume that both types + # hold equity; warn if the bisection on 𝜃1 ended at one of its bounds + if 𝜃1 > 1 - 0.002 or 𝜃1 < 0.3 + 0.002: + print(f'Warning: 𝜃1 = {𝜃1:.4f} is at a bisection bound, so both ' + 'types do not hold equity and the computed (k, b) does not ' + "satisfy the firm's first-order conditions; it is not an " + 'equilibrium of the special case studied here.') # Print the results print('finished') - # print('k,b,p,q,kfoc,bfoc,epstar,V,V_crit') - #formattedList = ["%.3f" % member for member in [kss, - # bss, - # pss, - # qss, - # kfoc, - # bfoc, - # epstar, - # Vss, - # V_crit]] - #print(formattedList) return kss,bss,Vss,qss,pss,c10ss,c11ss,c20ss,c21ss,𝜃1ss - #************************************************************* # Function: Equity and bond valuations by different agents #************************************************************* @@ -1109,7 +1007,8 @@ class BCG_incomplete_markets: Y = lambda 𝜖: np.exp(𝜖)*fk # Compute integration threshold - epstar = np.log(b/fk) + # Default threshold, kept inside the integration range [-bound, bound] + epstar = min(max(np.log(b/fk), -bound), bound) # Compute equity valuation with agent 1's IMRS intQ1 = lambda 𝜖: IMRS1(𝜖)*(Y(𝜖) - b) @@ -1129,7 +1028,6 @@ class BCG_incomplete_markets: return Q1,Q2,P1,P2 - #************************************************************* # Function: equilibrium valuations for firm, equity, bond #************************************************************* @@ -1194,12 +1092,11 @@ class BCG_incomplete_markets: ## Examples -Below we show some examples computed with the class `BCG_incomplete markets`. +Below we show some examples computed with the class `BCG_incomplete_markets`. ### First example -In the first example, we set up an instance of the BCG incomplete -markets model with default parameter values. +In the first example, we set up an instance of the BCG incomplete markets model with default parameter values. ```{code-cell} ipython3 :tags: [hide-output] @@ -1211,35 +1108,36 @@ kss,bss,Vss,qss,pss,c10ss,c11ss,c20ss,c21ss,𝜃1ss = mdl.solve_eq(print_crit=Fa ```{code-cell} ipython3 print(-kss+qss+pss*bss) print(Vss) +print(kss) +print(bss) print(𝜃1ss) ``` -Python reports to us that the equilibrium firm value is $V=0.101$, -with capital $k = 0.151$ and debt $b=0.484$. +Python reports to us that the equilibrium firm value is $V=0.101$, with capital $k = 0.151$ and debt $b=0.484$. -Let’s verify some things that have to be true if our algorithm has truly -found an equilibrium. +Let’s verify some things that have to be true if our algorithm has truly found an equilibrium. -Thus, let’s see if the firm is actually maximizing its firm value given -the equilibrium pricing function $q(k,b)$ for equity and -$p(k,b)$ for bonds. +Thus, let’s see if the firm is actually maximizing its firm value given the equilibrium pricing function $q(k,b)$ for equity and $p(k,b)$ for bonds. ```{code-cell} ipython3 kgrid, bgrid, Vgrid, Qgrid, Pgrid = mdl.eq_valuation(c10ss, c11ss, c20ss, c21ss,N=30) -print('Maximum valuation of the firm value in the (k,B) grid: {:.5f}'.format(Vgrid.max())) -print('Equilibrium firm value: {:.5f}'.format(Vss)) +i = np.unravel_index(np.argmax(Vgrid), Vgrid.shape) +print('Maximum firm value on the (k,b) grid: {:.5f} at k = {:.4f}, b = {:.4f}' + .format(Vgrid.max(), kgrid[i], bgrid[i])) +print('Firm value -K + q + p B at the equilibrium: {:.5f}' + .format(-kss + qss + pss * bss)) ``` -Up to the approximation involved in using a discrete grid, these numbers -give us comfort that the firm does indeed seem to be maximizing its -value at the top of the value hill on the $(k,b)$ plane that it -faces. +The grid maximum equals the firm value at the equilibrium to five digits, and it occurs at a value of $k$ close to $K$. + +Its location in $b$ is not informative: as the plots below show, firm value is almost flat in $b$ along a ridge, so many choices of $b$ attain nearly the same value. + +Up to the approximation involved in using a discrete grid, these numbers give us comfort that the firm is maximizing its value at the top of the value hill on the $(k,b)$ plane that it faces. Below we will plot the firm’s value as a function of $k,b$. -We’ll also plot the equilibrium price functions $q(k,b)$ and -$p(k,b)$. +We’ll also plot the equilibrium price functions $q(k,b)$ and $p(k,b)$. ```{code-cell} ipython3 from IPython.display import Image @@ -1277,77 +1175,55 @@ Image(fig.to_image(format="png", engine="kaleido")) #### A Modigliani-Miller theorem? -The red dot in the above graph is **both** an equilibrium $(b,k)$ -chosen by a representative firm **and** the equilibrium $B, K$ -pair chosen by the aggregate of all firms. +The red dot in the above graph is **both** an equilibrium $(k,b)$ chosen by a representative firm **and** the equilibrium $K, B$ pair chosen by the aggregate of all firms. -Thus, **in equilibrium** it -is true that +Thus, **in equilibrium** it is true that $$ -(b,k) = (B,K) +(k,b) = (K,B) $$ -But an individual firm named $\zeta \in [0,1]$ neither knows nor -cares whether it sets $(b(\zeta),k(\zeta)) = (B,K)$. +But an individual firm named $\zeta \in [0,1]$ neither knows nor cares whether it sets $(k(\zeta),b(\zeta)) = (K,B)$. -Indeed the above graph has a ridge of $b(\zeta)$’s that also -maximize the firm’s value so long as it sets $k(\zeta) = K$. +Indeed the above graph has a ridge of $b(\zeta)$’s that also maximize the firm’s value so long as it sets $k(\zeta) = K$. -Here it is important that the measure of firms that deviate from setting -$b$ at the red dot is very small – measure zero – so that -$B$ remains at the red dot even while one firm $\zeta$ -deviates. +Here it is important that the measure of firms that deviate from setting $b$ at the red dot is very small – measure zero – so that $B$ remains at the red dot even while one firm $\zeta$ deviates. -So within this equilibrium, there is a *qualified* Modigliani-Miller theorem -that asserts that firm $\zeta$’s value is -independent of how it mixes its financing between equity and bonds (so -long as it is not what other firms do on average). +So within this equilibrium, there is a *qualified* Modigliani-Miller theorem that asserts that firm $\zeta$’s value is independent of how it mixes its financing between equity and bonds (so long as other firms, on average, choose the equilibrium mix $B$ and the firm chooses $k = K$). -Thus, while an individual firm $\zeta$’s financial structure is -indeterminate, the **market’s** financial structure is determinant and -sits at the red dot in the above graph. +Thus, while an individual firm $\zeta$’s financial structure is indeterminate, the **market’s** financial structure is determinate and sits at the red dot in the above graph. -This contrasts sharply with the *unqualified* Modigliani-Miller theorem -descibed in the complete markets model in the lecture {doc}`Irrelevance of Capital Structure with Complete Markets `. +This contrasts sharply with the *unqualified* Modigliani-Miller theorem described in the complete markets model in the lecture {doc}`Irrelevance of Capital Structures with Complete Markets `. There the **market’s** financial structure was indeterminate. -These subtle distinctions bear more thought and exploration. +These subtle distinctions bear more thought and exploration. -So we will do some calculations to ferret out a sense in which -the equilibrium $(k,b) = (K,B)$ outcome at the red dot in the -above graph is **stable**. +So we will do some calculations to check whether nearby capital structures could also be equilibria, and in what sense the equilibrium $(k,b) = (K,B)$ outcome at the red dot in the above graph is isolated. -In particular, we’ll explore the consequences of some choices of -$b=B$ that deviate from the red dot and ask whether firm -$\zeta$ would want to remain at that $b$. +In particular, we’ll explore the consequences of some choices of $b=B$ that deviate from the red dot and ask whether firm $\zeta$ would want to remain at that $b$. In more detail, here is what we’ll do: 1. Obtain equilibrium values of capital and debt as $k^*=K$ and - $b^*=B$, the red dot above. -1. Now fix $k^*$ and let $b^{**} = b^* - e$ for some - $e > 0$. Conjecture that big $K = k^*$ but big + $b^*=B$, the red dot above. +1. Now fix $k^*$ and let $b^{**} = b^* + e$ for some + $e \neq 0$. Conjecture that big $K = k^*$ but big $B = b^{**}$. -1. Take $K$ and $B$ and compute intertermporal marginal rates of substitution (IMRS's) as we did before. -1. Taking the **new** IMRS to the firm’s problem. Plot 3D surface for +1. Take $K$ and $B$ and compute intertemporal marginal rates of substitution (IMRS's) as we did before. +1. Use the **new** IMRS in the firm’s problem and plot the 3D surface for the valuations of the firm with this **new** IMRS. 1. Check if the value at $k^*$, $b^{**}$ is at the top of this new 3D surface. -1. Repeat these calculations for $b^{**} = b^* + e$. +1. Repeat these calculations for a perturbation $e$ of the opposite sign. -To conduct the above procedures, we create a function `off_eq_check` -that inputs the BCG model instance parameters, equilibrium capital -$K=k^*$ and debt $B=b^*$, and a perturbation of debt $e$. +To conduct the above procedures, we create a function `off_eq_check` that inputs the BCG model instance parameters, equilibrium capital $K=k^*$ and debt $B=b^*$, and a perturbation of debt $e$. -The function outputs the fixed point firm values $V^{**}$, prices -$q^{**}$, $p^{**}$, and consumption choices $c^{**}$. +The function outputs the fixed point firm values $V^{**}$, prices $q^{**}$, $p^{**}$, and consumption choices $c^{**}$. Importantly, we relax the condition that only agent 2 holds bonds. -Now **both** agents can hold bonds, i.e., $0\leq \xi^1 \leq B$ and -$\xi^1 +\xi^2 = B$. +Now **both** agents can hold bonds, i.e., $0\leq \xi^1 \leq B$ and $\xi^1 +\xi^2 = B$. That implies the consumers’ budget constraints are: @@ -1355,12 +1231,12 @@ $$ \begin{aligned} c^1_0 &= w^1_0 + \theta^1_0V - q\theta^1 - p\xi^1 \\ c^2_0 &= w^2_0 + \theta^2_0V - q\theta^2 - p\xi^2 \\ -c^1_1(\epsilon) &= w^1_1(\epsilon) + \theta^1 d^e(k,b;\epsilon) + \xi^1 \\ -c^2_1(\epsilon) &= w^2_1(\epsilon) + \theta^2 d^e(k,b;\epsilon) + \xi^2 +c^1_1(\epsilon) &= w^1_1(\epsilon) + \theta^1 d^e(k,b;\epsilon) + \xi^1 d^b(k,b;\epsilon) \\ +c^2_1(\epsilon) &= w^2_1(\epsilon) + \theta^2 d^e(k,b;\epsilon) + \xi^2 d^b(k,b;\epsilon) \end{aligned} $$ -The function also outputs agent 1’s bond holdings $\xi_1$. +The function also outputs agent 1’s bond holdings $\xi^1$. ```{code-cell} ipython3 def off_eq_check(mdl,kss,bss,e=0.1): @@ -1395,8 +1271,7 @@ def off_eq_check(mdl,kss,bss,e=0.1): intpp1b = njit(lambda 𝜖, fk, 𝜃1, 𝜓1, 𝜉1, b: (w11(𝜖) + 𝜃1*(Y(𝜖, fk)-b) + 𝜉1)**(-𝜓1)*g(𝜖)) intpp2a = njit(lambda 𝜖, fk, 𝜓2, 𝜉2, b: (Y(𝜖, fk)/b)*(w21(𝜖) + Y(𝜖, fk)/b*𝜉2)**(-𝜓2)*g(𝜖)) intpp2b = njit(lambda 𝜖, fk, 𝜃2, 𝜓2, 𝜉2, b: (w21(𝜖) + 𝜃2*(Y(𝜖, fk)-b) + 𝜉2)**(-𝜓2)*g(𝜖)) - intqq2 = njit(lambda 𝜖, fk, 𝜃2, 𝜓2, b: (w21(𝜖) + 𝜃2*(Y(𝜖, fk)-b) + b)**(-𝜓2)*(Y(𝜖, fk) - b)*g(𝜖)) - + intqq2 = njit(lambda 𝜖, fk, 𝜃2, 𝜓2, 𝜉2, b: (w21(𝜖) + 𝜃2*(Y(𝜖, fk)-b) + 𝜉2)**(-𝜓2)*(Y(𝜖, fk) - b)*g(𝜖)) # Loop: Find fixed points V, q and p V_crit = 1 @@ -1409,11 +1284,10 @@ def off_eq_check(mdl,kss,bss,e=0.1): # Production fk = A*(k**𝛼) -# Y = lambda 𝜖: np.exp(𝜖)*fk # Compute integration threshold - epstar = np.log(b/fk) - + # Default threshold, kept inside the integration range [-bound, bound] + epstar = min(max(np.log(b/fk), -bound), bound) #************************************************************** # Compute the prices and allocations consistent with consumers' @@ -1421,8 +1295,7 @@ def off_eq_check(mdl,kss,bss,e=0.1): #************************************************************** # We impose the following: - # Agent 1 buys equity - # Agent 2 buys equity and all debt + # Both agents may buy equity and debt # Agents trade such that prices converge #======== @@ -1448,8 +1321,6 @@ def off_eq_check(mdl,kss,bss,e=0.1): ## First, compute the constant term that is not influenced by q ## that is, 𝛽E[u'(c^{1}_{1})d^{e}(k,B)] -# intqq1 = lambda 𝜖: (w11(𝜖) + 𝜃1*(Y(𝜖, fk) - b) + 𝜉1)**(-𝜓1)*(Y(𝜖, fk) - b)*g(𝜖) -# const_qq1 = 𝛽 * quad(intqq1,epstar,bound)[0] const_qq1 = 𝛽 * quad(intqq1,epstar,bound, args=(fk, 𝜃1, 𝜓1, 𝜉1, b))[0] ## Second, iterate to get the equity price q @@ -1465,14 +1336,11 @@ def off_eq_check(mdl,kss,bss,e=0.1): qq1h = qq1 diff = abs(qq1l-qq1h) - # pp1 is the bond price consistent with agent-2 Euler Equation + # pp1 is the bond price consistent with agent-1 Euler Equation ## Note: Price is in the date-0 budget constraint of the agent ## First, compute the constant term that is not influenced by p ## that is, 𝛽E[u'(c^{1}_{1})d^{b}(k,B)] -# intpp1a = lambda 𝜖: (Y(𝜖, fk)/b)*(w11(𝜖) + Y(𝜖, fk)/b*𝜉1)**(-𝜓1)*g(𝜖) -# intpp1b = lambda 𝜖: (w11(𝜖) + 𝜃1*(Y(𝜖, fk)-b) + 𝜉1)**(-𝜓1)*g(𝜖) -# const_pp1 = 𝛽 * (quad(intpp1a,-bound,epstar)[0] + quad(intpp1b,epstar,bound)[0]) const_pp1 = 𝛽 * (quad(intpp1a,-bound,epstar, args=(fk, 𝜓1, 𝜉1, b))[0] \ + quad(intpp1b,epstar,bound, args=(fk, 𝜃1, 𝜓1, 𝜉1, b))[0]) @@ -1500,9 +1368,6 @@ def off_eq_check(mdl,kss,bss,e=0.1): ## First, compute the constant term that is not influenced by p ## that is, 𝛽E[u'(c^{2}_{1})d^{b}(k,B)] -# intpp2a = lambda 𝜖: (Y(𝜖, fk)/b)*(w21(𝜖) + Y(𝜖, fk)/b*𝜉2)**(-𝜓2)*g(𝜖) -# intpp2b = lambda 𝜖: (w21(𝜖) + 𝜃2*(Y(𝜖, fk)-b) + 𝜉2)**(-𝜓2)*g(𝜖) -# const_pp2 = 𝛽 * (quad(intpp2a,-bound,epstar)[0] + quad(intpp2b,epstar,bound)[0]) const_pp2 = 𝛽 * (quad(intpp2a,-bound,epstar, args=(fk, 𝜓2, 𝜉2, b))[0] \ + quad(intpp2b,epstar,bound, args=(fk, 𝜃2, 𝜓2, 𝜉2, b))[0]) @@ -1520,14 +1385,11 @@ def off_eq_check(mdl,kss,bss,e=0.1): diff = abs(pp2l-pp2h) # p be the maximum valuation for the bond among agents - ## This will be the equity price based on Makowski's criterion + ## This will be the bond price based on Makowski's criterion p = max(pp1,pp2) - # qq2 is the equity price consistent with agent-2 Euler Equation -# intqq2 = lambda 𝜖: (w21(𝜖) + 𝜃2*(Y(𝜖, fk)-b) + b)**(-𝜓2)*(Y(𝜖, fk) - b)*g(𝜖) -# const_qq2 = 𝛽 * quad(intqq2,epstar,bound)[0] - const_qq2 = 𝛽 * quad(intqq2,epstar,bound, args=(fk, 𝜃2, 𝜓2, b))[0] + const_qq2 = 𝛽 * quad(intqq2,epstar,bound, args=(fk, 𝜃2, 𝜓2, 𝜉2, b))[0] qq2l = 0 qq2h = ww20 diff = 1 @@ -1552,14 +1414,11 @@ def off_eq_check(mdl,kss,bss,e=0.1): else: 𝜃1b = 𝜃1 - #print(p,q,𝜉1,𝜃1) - if pp1 > pp2: 𝜉1a = 𝜉1 else: 𝜉1b = 𝜉1 - #================ # Get consumption #================ @@ -1581,22 +1440,15 @@ def off_eq_check(mdl,kss,bss,e=0.1): Here is our strategy for checking *stability* of an equilibrium. -We use `off_eq_check` to obtain consumption plans for both agents -at the conjectured big $K$ and big $B$. +We use `off_eq_check` to obtain consumption plans for both agents at the conjectured big $K$ and big $B$. -Then we input consumption plans into the function `eq_valuation` -from the BCG model class and plot the agents’ valuations associated -with different choices of $k$ and $b$. +Then we input consumption plans into the function `eq_valuation` from the BCG model class and plot the agents’ valuations associated with different choices of $k$ and $b$. -Our hunch is that $(k^*,b^{**})$ is **not** at the top of the -firm valuation 3D surface so that the firm is **not** maximizing its -value if it chooses $k = K = k^*$ and $b = B = b^{**}$. +Our hunch is that $(k^*,b^{**})$ is **not** at the top of the firm valuation 3D surface so that the firm is **not** maximizing its value if it chooses $k = K = k^*$ and $b = B = b^{**}$. -That indicates that $(k^*,b^{**})$ is not an equilibrium capital -structure for the firm. +That indicates that $(k^*,b^{**})$ is not an equilibrium capital structure for the firm. -We first check the case in which $b^{**} = b^* - e$ where -$e = 0.1$: +We first check the case in which $b^{**} = b^* + e$ where $e = -0.1$: ```{code-cell} ipython3 #====================== Experiment 1 ======================# @@ -1632,23 +1484,19 @@ fig.update_layout(scene = dict( fig.update_layout(scene_camera=dict(eye=dict(x=1.5, y=-1.5, z=2))) fig.update_layout(title='Equilibrium firm valuation for the grid of (k,b)') - # Export to PNG file Image(fig.to_image(format="png", engine="kaleido")) # fig.show() will provide interactive plot when running # code locally ``` -In the above 3D surface of prospective firm valuations, the perturbed -choice $(k^*,b^{*}-e)$, represented by the red dot, is not at the -top. +In the above 3D surface of prospective firm valuations, the perturbed choice $(k^*,b^{*}-0.1)$, represented by the red dot, is not at the top. -The firm could issue more debts and attain a higher firm valuation from -the market. +The firm could issue more debt and attain a higher firm valuation from the market. -Therefore, $(k^*,b^{*}-e)$ would not be an equilibrium. +Therefore, $(k^*,b^{*}-0.1)$ would not be an equilibrium. -Next, we check for $b^{**} = b^* + e$. +Next, we check for $b^{**} = b^* + e$ where $e = 0.1$. ```{code-cell} ipython3 #====================== Experiment 2 ======================# @@ -1690,36 +1538,29 @@ Image(fig.to_image(format="png", engine="kaleido")) # code locally ``` -In contrast to $(k^*,b^* - e)$, the 3D surface for -$(k^*,b^*+e)$ now indicates that a firm would want o *decrease* -its debt issuance to attain a higher valuation. +In contrast to $(k^*,b^* - 0.1)$, the 3D surface for $(k^*,b^*+0.1)$ now indicates that a firm would want to *decrease* its debt issuance to attain a higher valuation. -That incentive to deviate means that $(k^*,b^*+e)$ is not an -equilibrium capital structure for the firm. +That incentive to deviate means that $(k^*,b^*+0.1)$ is not an equilibrium capital structure for the firm. -Interestingly, if consumers were to anticipate that firms would -over-issue debt, i.e. $B > b^*$, then both types of consumer would -want to hold corporate debt. +Interestingly, if consumers were to anticipate that firms would over-issue debt, i.e. $B > b^*$, then both types of consumer would want to hold corporate debt. -For example, $\xi^1 > 0$: +Type 2 consumers would then value equity less than type 1 consumers do, so they would hold no equity; this perturbed economy therefore lies outside the special case in which both types hold equity. + +For example, $\xi^1 > 0$: ```{code-cell} ipython3 print('Bond holdings of agent 1: {:.3f}'.format(𝜉1e2)) ``` -Our two *stability experiments* suggest that the equilibrium capital -structure $(k^*,b^*)$ is locally unique even though **at the -equilibrium** an individual firm would be willing to deviate from the -representative firms’ equilibrium debt choice. +Our two experiments show that when the representative firms' debt is $B = b^* \pm 0.1$, an individual firm would want to move its debt toward $b^*$, and in fact past it, so neither perturbed value of $B$ is an equilibrium. + +This is consistent with $(k^*,b^*)$ being an isolated equilibrium, although the experiments establish neither its uniqueness nor its dynamic stability, and it holds even though **at the equilibrium** an individual firm would be willing to deviate from the representative firms’ equilibrium debt choice. -These experiments thus refine our discussion of the *qualified* -Modigliani-Miller theorem that prevails in this example economy. +These experiments thus refine our discussion of the *qualified* Modigliani-Miller theorem that prevails in this example economy. #### Equilibrium equity and bond price functions -It is also interesting to look at the equilibrium price functions -$q(k,b)$ and $p(k,b)$ faced by firms in our rational -expectations equilibrium. +It is also interesting to look at the equilibrium price functions $q(k,b)$ and $p(k,b)$ faced by firms in our rational expectations equilibrium. ```{code-cell} ipython3 # Equity Valuation @@ -1744,7 +1585,6 @@ fig.update_layout(scene = dict( fig.update_layout(scene_camera=dict(eye=dict(x=1.5, y=-1.5, z=2))) fig.update_layout(title='Equilibrium equity valuation for the grid of (k,b)') - # Export to PNG file Image(fig.to_image(format="png", engine="kaleido")) # fig.show() will provide interactive plot when running @@ -1766,7 +1606,7 @@ fig = go.Figure(data=[go.Scatter3d(x=[kss], fig.update_layout(scene = dict( xaxis_title='x - Capital k', yaxis_title='y - Debt b', - zaxis_title='z - Bond price q', + zaxis_title='z - Bond price p', aspectratio = dict(x=1,y=1,z=1)), width=700, height=700, @@ -1774,7 +1614,6 @@ fig.update_layout(scene = dict( fig.update_layout(scene_camera=dict(eye=dict(x=1.5, y=-1.5, z=2))) fig.update_layout(title='Equilibrium bond valuation for the grid of (k,b)') - # Export to PNG file Image(fig.to_image(format="png", engine="kaleido")) # fig.show() will provide interactive plot when running @@ -1783,22 +1622,15 @@ Image(fig.to_image(format="png", engine="kaleido")) ### Comments on equilibrium pricing functions -The equilibrium pricing functions displayed above merit study and -reflection. +The equilibrium pricing functions displayed above merit study and reflection. -They reveal the countervailing effects on a firm’s valuations of bonds -and equities that lie beneath the Modigliani-Miller ridge apparent in -our earlier graph of an individual firm $\zeta$’s value as a -function of $k(\zeta), b(\zeta)$. +They reveal the countervailing effects on a firm’s valuations of bonds and equities that lie beneath the Modigliani-Miller ridge apparent in our earlier graph of an individual firm $\zeta$’s value as a function of $k(\zeta), b(\zeta)$. ### Another example economy -We illustrate how the fraction of initial endowments held by agent 2, -$w^2_0/(w^1_0+w^2_0)$ affects an equilibrium capital structure -$(k,b) = (K, B)$ well as associated equilibrium allocations. +We illustrate how the fraction of initial endowments held by agent 2, $w^2_0/(w^1_0+w^2_0)$ affects an equilibrium capital structure $(k,b) = (K, B)$ as well as associated equilibrium allocations. -We are interested in how agents 1 and 2 -value equity and bond. +We are interested in how agents 1 and 2 value equity and bonds. $$ \begin{aligned} @@ -1807,8 +1639,7 @@ P^i = \beta \int \frac{u^\prime(C^{i,*}_1(\epsilon))}{u^\prime(C^{i,*}_0)} d^b(k \end{aligned} $$ -The function `valuations_by_agent` is used in calculating these -valuations. +The function `valuations_by_agent` is used in calculating these valuations. ```{code-cell} ipython3 :tags: [hide-output] @@ -1895,14 +1726,9 @@ plt.show() Please stare at the above panels. -They describe how equilibrium prices and quantities respond to -alterations in the structure of society’s *hedging desires* across -economies with different allocations of the initial endowment to our two -types of agents. +They describe how equilibrium prices and quantities respond to alterations in the structure of society’s *hedging desires* across economies with different allocations of the initial endowment to our two types of agents. -Now let’s see how the two types of agents value bonds and equities, -keeping in mind that the type that values the asset highest determines -the equilibrium price (and thus the pertinent set of Big $C$’s). +Now let’s see how the two types of agents value bonds and equities, keeping in mind that the type that values the asset highest determines the equilibrium price (and thus the pertinent set of Big $C$’s). ```{code-cell} ipython3 # Comparing the prices @@ -1931,10 +1757,315 @@ plt.show() It is rewarding to stare at the above plots too. -In equilibrium, equity valuations are the same across the two types of -agents but bond valuations are not. +In equilibrium, equity valuations are the same across the two types of agents but bond valuations are not. Agents of type 2 value bonds more highly (they want more hedging). -Taken together with our earlier plot of equity holdings, these graphs confirm our earlier conjecture that while both type -of agents hold equities, only agents of type 2 holds bonds. +Taken together with our earlier plot of equity holdings, these graphs confirm our earlier conjecture that while both types of agents hold equities, only agents of type 2 hold bonds. + +## Exercises + +```{exercise} +:label: bcgi_ex1 + +Type 2 consumers hold all of the firm's bonds because their period $1$ endowment $w_1^2(\epsilon)$ loads heavily on the productivity shock $\epsilon$ through the parameter $\chi_2$. + +In this exercise we vary the strength of that hedging motive. + +Holding all other parameters at their default values (but setting `ktop = 0.5` and `btop = 2.5` so that the bisection brackets are wide), solve for an equilibrium for each $\chi_2 \in \{0.5, 0.6, 0.7, 0.8, 0.9\}$. + +For each value of $\chi_2$ report + +* equilibrium capital $k$ and debt $b$ +* the default threshold $\epsilon^* = \log\left(b / (A k^\alpha)\right)$ and the probability of default $\textrm{Prob}(\epsilon < \epsilon^*)$ +* the bond price $p$ and the two agents' valuations $P^1, P^2$ of a bond +* agent 1's equity share $\theta^1$ + +Before you compute anything, predict how leverage $b$ and the probability of default respond to an increase in $\chi_2$, and explain your prediction. + +Then check whether the computed equilibria stay within the "special case" assumed in the lecture, namely $0 < \theta^1 < 1$ and $P^1 < P^2$. +``` + +```{solution-start} bcgi_ex1 +:class: dropdown +``` + +The method `solve_eq` prints `finished` when it is done, so we wrap it in a helper that suppresses that message but passes on any warning. + +```{code-cell} ipython3 +import io +import contextlib +from scipy.stats import norm + +def solve_quietly(**kwargs): + "Solve for an equilibrium, suppressing the 'finished' message but not warnings." + model = BCG_incomplete_markets(ktop=0.5, btop=2.5, **kwargs) + buffer = io.StringIO() + with contextlib.redirect_stdout(buffer): + results = model.solve_eq(print_crit=False) + for line in buffer.getvalue().splitlines(): + if line.startswith('Warning'): + print(line) + return model, results + +def summarize(model, results): + "Compute default threshold, default probability and agents' valuations." + k, b, V, q, p, c10, c11, c20, c21, 𝜃1 = results + eps_star = np.log(b / (model.A * k**model.𝛼)) + prob_default = norm.cdf(eps_star, loc=model.𝜇, scale=model.𝜎) + Q1, Q2, P1, P2 = model.valuations_by_agent(c10, c11, c20, c21, k, b) + return dict(k=k, b=b, V=V, q=q, p=p, theta1=𝜃1, eps_star=eps_star, + prob_default=prob_default, Q1=Q1, Q2=Q2, P1=P1, P2=P2) +``` + +```{code-cell} ipython3 +𝜒2_values = [0.5, 0.6, 0.7, 0.8, 0.9] +ex1 = [] +for 𝜒2 in 𝜒2_values: + model, results = solve_quietly(𝜒2=𝜒2) + ex1.append(summarize(model, results)) + +print(" 𝜒2 k b eps* P(def) p P1 P2 𝜃1") +for 𝜒2, r in zip(𝜒2_values, ex1): + print(f"{𝜒2:.1f} {r['k']:.4f} {r['b']:.4f} {r['eps_star']:7.3f} " + f"{r['prob_default']:.3f} {r['p']:.4f} {r['P1']:.4f} " + f"{r['P2']:.4f} {r['theta1']:.3f}") +``` + +```{code-cell} ipython3 +fig, ax = plt.subplots(2, 2, figsize=(10, 7)) +panels = [('b', 'debt $b$'), ('k', 'capital $k$'), + ('eps_star', r'default threshold $\epsilon^*$'), + ('prob_default', 'probability of default')] +for axis, (key, title) in zip(ax.flatten(), panels): + axis.plot(𝜒2_values, [r[key] for r in ex1], marker='o') + axis.set_title(title) + axis.set_xlabel(r'$\chi_2$') +plt.tight_layout() +plt.show() +``` + +As $\chi_2$ rises from $0.5$ to $0.9$, equilibrium debt rises from about $0.389$ to $0.484$, capital rises from about $0.136$ to $0.151$, the default threshold rises from about $-0.66$ to $-0.51$, and the probability of default roughly doubles, from about $5.6\%$ to $11.4\%$. + +The economics runs through type 2 consumers' hedging demand. + +A larger $\chi_2$ makes a type 2 consumer's period $1$ endowment more exposed to the productivity shock. + +A bond pays a constant amount except in low-$\epsilon$ default states, so it is a better hedge than equity for a consumer whose endowment is already high when $\epsilon$ is high. + +Type 2 consumers therefore value bonds more highly: their valuation $P^2 = p$ rises from about $0.346$ to $0.376$, while type 1 consumers' valuation $P^1$ barely moves and stays below $p$. + +Firms respond to the higher price that the marginal bondholder is willing to pay by issuing more bonds. + +Because debt rises faster than expected output $A k^\alpha$, the default threshold $\epsilon^*$ rises, so the extra debt is riskier. + +Type 1 consumers absorb more of the equity: $\theta^1$ rises from about $0.77$ to $0.99$. + +For every value of $\chi_2$ we have $0 < \theta^1 < 1$ and $P^1 < P^2$, so the computed equilibria stay within the special case under which the lecture derives the firm's first-order conditions. + +Notice, however, that at the default value $\chi_2 = 0.9$ agent 1 already holds about $99\%$ of the equity, close to the corner $\theta^1 = 1$ at which that special case breaks down. + +```{solution-end} +``` + +```{exercise} +:label: bcgi_ex2 + +Now study how the riskiness of production affects capital structure. + +To isolate the effect of *risk*, hold the mean of the productivity factor fixed at its default value $E\left[e^\epsilon\right] = e^{\mu + \sigma^2/2} = e^{0.055}$ by setting $\mu = 0.055 - \sigma^2/2$, and solve for an equilibrium for each $\sigma \in \{0.40, 0.45, 0.50, 0.55, 0.60\}$. + +(Because the endowment functions $w_1^i(\epsilon)$ are normalized to have mean one, this change is a mean-preserving spread of both output and endowments.) + +For each $\sigma$ report $k$, $b$, $\epsilon^*$, the probability of default, the prices $q$ and $p$, and $\theta^1$. + +Explain why the default threshold and the probability of default can move in *opposite* directions. +``` + +```{solution-start} bcgi_ex2 +:class: dropdown +``` + +We reuse the helpers `solve_quietly` and `summarize` from the solution to {ref}`bcgi_ex1`. + +```{code-cell} ipython3 +𝜎_values = [0.40, 0.45, 0.50, 0.55, 0.60] +ex2 = [] +for 𝜎 in 𝜎_values: + 𝜇 = 0.055 - 𝜎**2 / 2 + model, results = solve_quietly(𝜎=𝜎, 𝜇=𝜇) + ex2.append(summarize(model, results)) + +print(" 𝜎 k b eps* P(def) q p 𝜃1") +for 𝜎, r in zip(𝜎_values, ex2): + print(f"{𝜎:.2f} {r['k']:.4f} {r['b']:.4f} {r['eps_star']:7.3f} " + f"{r['prob_default']:.3f} {r['q']:.4f} {r['p']:.4f} {r['theta1']:.3f}") +``` + +```{code-cell} ipython3 +fig, ax = plt.subplots(1, 3, figsize=(13, 4)) +ax[0].plot(𝜎_values, [r['b'] for r in ex2], marker='o', label='debt $b$') +ax[0].plot(𝜎_values, [r['k'] for r in ex2], marker='o', label='capital $k$') +ax[0].legend() +ax[1].plot(𝜎_values, [r['eps_star'] for r in ex2], marker='o') +ax[1].set_title(r'default threshold $\epsilon^*$') +ax[2].plot(𝜎_values, [r['prob_default'] for r in ex2], marker='o') +ax[2].set_title('probability of default') +for axis in ax: + axis.set_xlabel(r'$\sigma$') +plt.tight_layout() +plt.show() +``` + +As $\sigma$ rises from $0.40$ to $0.60$ (holding $E[e^\epsilon]$ fixed), capital rises from about $0.151$ to $0.185$, while debt edges *down* from about $0.484$ to $0.475$. + +The default threshold falls from about $-0.51$ to $-0.65$, yet the probability of default *rises* from about $11.4\%$ to $19.1\%$. + +The bond price rises sharply, from about $0.376$ to $0.511$, while the equity price falls slightly, from about $0.070$ to $0.066$, and agent 1's equity share falls from about $0.986$ to $0.967$. + +With $\gamma = 3$ marginal utility is convex, so a mean-preserving spread in period $1$ resources strengthens consumers' precautionary motive to transfer resources to period $1$. + +That raises the value that the marginal investor attaches to period $1$ payoffs, which shows up in a higher bond price and, through the first-order condition {eq}`Eqn1`, in a higher level of investment. + +Because output $A k^\alpha$ rises while debt barely changes, the ratio $b/(A k^\alpha)$ falls and so does the threshold $\epsilon^*$. + +The default threshold $\epsilon^*$ is a point on the $\epsilon$ axis, while the probability of default is the mass that the density $g$ puts below that point. + +A mean-preserving spread thickens the left tail of $g$, so the probability of default can rise even when $\epsilon^*$ falls. + +```{solution-end} +``` + +```{exercise} +:label: bcgi_ex3 + +This exercise dissects the *qualified* Modigliani-Miller result of the lecture. + +Solve for the equilibrium at the default parameter values and hold the equilibrium consumption plans $C^i_0, C^i_1(\epsilon)$ fixed. + +Fix capital at its equilibrium value $k = K$, and for a grid of 71 values of $b$ in $[0.1, 0.8]$ use `valuations_by_agent` to compute each agent's valuations $Q^i(K,b)$ and $P^i(K,b)$ of equity and bonds. + +1. Compute the firm value $V(K,b) = -K + \max_i Q^i(K,b) + b \max_i P^i(K,b)$, and show that it is (numerically) independent of $b$, even though $q(K,b)$, $p(K,b) b$, and the probability of default all vary with $b$. +1. Show that $Q^2(K,b) + b P^2(K,b) = \beta \int \frac{u'(C_1^2(\epsilon))}{u'(C_0^2)} A K^\alpha e^\epsilon g(\epsilon) d\epsilon$ for every $b$, and explain why this identity implies the flat ridge. +1. Plot $D(b) = Q^1(K,b) - Q^2(K,b)$. Show that $D(b) \leq 0$ on the grid, and that $D$ is maximized near the equilibrium debt level $B$. +1. Differentiate $D(b)$ and relate the condition $D'(B) = 0$ to the firm's first-order condition {eq}`Eqn2`. Use this to explain why the *aggregate* debt level $B$ is determinate even though an individual firm's debt level is not. +``` + +```{solution-start} bcgi_ex3 +:class: dropdown +``` + +The lecture's loop over initial endowments overwrote the baseline model, so we solve it again. + +```{code-cell} ipython3 +from scipy.integrate import quad + +mdl = BCG_incomplete_markets() +kss, bss, Vss, qss, pss, c10ss, c11ss, c20ss, c21ss, 𝜃1ss = mdl.solve_eq(print_crit=False) + +b_grid = np.linspace(0.1, 0.8, 71) +Q1, Q2, P1, P2 = np.array([mdl.valuations_by_agent(c10ss, c11ss, c20ss, c21ss, kss, b) + for b in b_grid]).T + +q_grid = np.maximum(Q1, Q2) +p_grid = np.maximum(P1, P2) +V_grid = -kss + q_grid + p_grid * b_grid + +fk = mdl.A * kss**mdl.𝛼 +eps_star = np.log(b_grid / fk) +prob_default = norm.cdf(eps_star, loc=mdl.𝜇, scale=mdl.𝜎) + +print(f"K = {kss:.4f}, B = {bss:.4f}") +print(f"range of V(K,b) over the grid: {V_grid.min():.8f} to {V_grid.max():.8f}") +print(f"range of q(K,b): {q_grid.min():.4f} to {q_grid.max():.4f}") +print(f"range of p(K,b) b: {(p_grid*b_grid).min():.4f} to {(p_grid*b_grid).max():.4f}") +print(f"range of default probability: {prob_default.min():.4f} to {prob_default.max():.4f}") +``` + +```{code-cell} ipython3 +# Agent 2's valuation of the firm's entire output A K^alpha e^epsilon +IMRS2 = lambda 𝜖: mdl.𝛽 * (c21ss(𝜖) / c20ss)**(-mdl.𝜓2) * mdl.g(𝜖) +whole_firm = quad(lambda 𝜖: IMRS2(𝜖) * fk * np.exp(𝜖), -mdl.bound, mdl.bound)[0] + +print(f"agent 2's value of output: {whole_firm:.8f}") +print(f"max |Q2 + b P2 - that value|: {np.max(np.abs(Q2 + b_grid*P2 - whole_firm)):.2e}") +print(f"max P1 - P2 on the grid: {np.max(P1 - P2):.4f}") +``` + +```{code-cell} ipython3 +D = Q1 - Q2 + +fig, ax = plt.subplots(1, 2, figsize=(12, 4)) +ax[0].plot(b_grid, q_grid, label='equity value $q(K,b)$') +ax[0].plot(b_grid, p_grid * b_grid, label='bond value $p(K,b)\\,b$') +ax[0].plot(b_grid, V_grid + kss, label='$q + p b = V + K$', linestyle='--') +ax[0].axvline(bss, color='gray', linestyle=':') +ax[0].set_xlabel('$b$') +ax[0].legend() + +ax[1].plot(b_grid, D) +ax[1].axvline(bss, color='gray', linestyle=':', label='equilibrium $B$') +ax[1].axhline(0, color='black', lw=0.5) +ax[1].set_xlabel('$b$') +ax[1].set_title('$D(b) = Q^1(K,b) - Q^2(K,b)$') +ax[1].legend() +plt.tight_layout() +plt.show() + +print(f"max D on grid: {D.max():.2e} at b = {b_grid[D.argmax()]:.3f}") +``` + +**Part 1.** Over $b \in [0.1, 0.8]$, the firm value $V(K,b)$ varies only between $0.10073832$ and $0.10073888$, a range of less than $10^{-6}$ that reflects quadrature and solver tolerances. + +Meanwhile the equity value $q(K,b)$ falls from about $0.21$ to $0.02$, the value of debt $p(K,b)\,b$ rises from about $0.04$ to $0.23$, and the probability of default rises from essentially $0$ to about $52\%$. + +A firm $\zeta$ that holds $k(\zeta) = K$ is thus indifferent among all of these capital structures: the left panel displays the Modigliani-Miller ridge. + +**Part 2.** Because $d^e(K,b;\epsilon) + b\, d^b(K,b;\epsilon) = \max\{A K^\alpha e^\epsilon - b, 0\} + \min\{A K^\alpha e^\epsilon, b\} = A K^\alpha e^\epsilon$ for every $\epsilon$, valuing both claims with the *same* stochastic discount factor gives + +$$ +Q^2(K,b) + b P^2(K,b) = \beta \int \frac{u'(C_1^2(\epsilon))}{u'(C_0^2)} A K^\alpha e^\epsilon g(\epsilon) d\epsilon , +$$ + +which does not depend on $b$. + +The computation confirms this: the two sides differ by at most $5.5 \times 10^{-7}$ on the grid. + +Since $P^1 < P^2$ at every $b$ on the grid (the largest value of $P^1 - P^2$ is about $-0.032$), $p(K,b) = P^2(K,b)$, and so + +$$ +V(K,b) = -K + Q^2(K,b) + b P^2(K,b) + \max\{D(b), 0\} . +$$ + +The first three terms are independent of $b$, so $V(K,b)$ is flat in $b$ exactly where $D(b) \leq 0$, that is, where type 2 consumers are (weakly) the marginal holders of *both* equity and bonds. + +This is the complete-markets logic of Modigliani and Miller operating locally: when a single stochastic discount factor prices every claim that the firm issues, splitting a given output stream into debt and equity cannot change its total value. + +**Part 3.** The right panel shows that $D(b) < 0$ at every grid point, with a maximum of about $-4.2 \times 10^{-5}$ at $b = 0.490$, next to the equilibrium value $B \approx 0.484$. + +(In an exact equilibrium $D(B) = 0$, since both types hold equity; the small negative value reflects the tolerance of $0.001$ used in the bisection on $\theta^1$.) + +Notice also that $D(b)$ is very flat to the right of $B$, where it stays within about $2 \times 10^{-4}$ of zero. + +**Part 4.** Differentiating $D(b)$ with the Leibniz rule, the boundary terms vanish because $d^e(K,b;\epsilon^*) = 0$, so + +$$ +D'(b) = -\beta \int_{\epsilon^*}^\infty \frac{u'(C_1^1(\epsilon))}{u'(C_0^1)} g(\epsilon) d\epsilon + \beta \int_{\epsilon^*}^\infty \frac{u'(C_1^2(\epsilon))}{u'(C_0^2)} g(\epsilon) d\epsilon . +$$ + +Setting $D'(B) = 0$ is precisely the firm's first-order condition {eq}`Eqn2` for debt. + +The two conditions $D(B) = 0$ and $D(b) \leq 0$ near $B$ require $B$ to be a local maximum of $D$, hence $D'(B) = 0$. + +The argument makes clear what is determinate and what is not. + +Given the aggregate consumption plans, an individual firm faces a flat ridge and does not care how it finances $K$. + +But the consumption plans $C^i$ themselves depend on the aggregate debt $B$ that consumers hold, and an equilibrium $B$ must generate stochastic discount factors that satisfy $D'(B) = 0$. + +If aggregate debt were at some other level, the two types' valuations of a marginal unit of equity would no longer be tangent at $B$, the ridge would disappear, and an individual firm could raise its value by changing its debt, as the lecture's two stability experiments with $B = b^* \pm 0.1$ illustrate. + +In the complete markets economy, by contrast, a single stochastic discount factor prices every claim, the counterpart of $D(b)$ is identically zero for every aggregate $B$, and so the aggregate capital structure is indeterminate too. + +```{solution-end} +``` diff --git a/lectures/_static/quant-econ.bib b/lectures/_static/quant-econ.bib index df2e3331..e83d20a9 100644 --- a/lectures/_static/quant-econ.bib +++ b/lectures/_static/quant-econ.bib @@ -536,6 +536,16 @@ @article{Modigliani_Miller_1958 pages = {261-297} } +@article{Makowski_1983, + author = {Louis Makowski}, + title = {Competition and Unanimity Revisited}, + journal = {American Economic Review}, + volume = {73}, + number = {3}, + year = {1983}, + pages = {329-339} +} + @article{definetti, Author = {Bruno de Finetti}, Date-Added = {2014-12-26 17:45:57 +0000}, @@ -896,6 +906,16 @@ @book{leamer1978specification publisher={John Wiley \& Sons Incorporated} } +@article{merton1980estimating, + title={On estimating the expected return on the market: An exploratory investigation}, + author={Merton, Robert C}, + journal={Journal of Financial Economics}, + volume={8}, + number={4}, + pages={323--361}, + year={1980} +} + @article{black1992global, title={Global portfolio optimization}, author={Black, Fischer and Litterman, Robert}, diff --git a/lectures/black_litterman.md b/lectures/black_litterman.md index 807c1ee9..012ed518 100644 --- a/lectures/black_litterman.md +++ b/lectures/black_litterman.md @@ -27,7 +27,7 @@ kernelspec: ## Overview -This lecture describes extensions to the classical mean-variance portfolio theory summarized in our lecture [Elementary Asset Pricing Theory](https://python-advanced.quantecon.org/asset_pricing_lph.html). +This lecture describes extensions to the classical mean-variance portfolio theory summarized in our lecture {doc}`Elementary Asset Pricing Theory `. The classic theory described there assumes that a decision maker completely trusts the statistical model that he posits to govern the joint distribution of returns on a list of available assets. @@ -35,25 +35,16 @@ Both extensions described here put distrust of that statistical model into the m One is a model of Black and Litterman {cite}`black1992global` that imputes to the decision maker distrust of historically estimated mean returns but still complete trust of estimated covariances of returns. -The second model also imputes to the decision maker doubts about his statistical model, but now by saying that, because of that distrust, the decision maker uses a version of robust control theory described in this lecture [Robustness](https://python-advanced.quantecon.org/robustness.html). +The second model also imputes to the decision maker doubts about his statistical model, but now by saying that, because of that distrust, the decision maker uses a version of robust control theory described in the lecture {doc}`Robustness `. -The famous **Black-Litterman** (1992) {cite}`black1992global` portfolio choice model was motivated by the finding that with high frequency or -moderately high frequency data, means are more difficult to estimate than -variances. +The famous **Black-Litterman** (1992) {cite}`black1992global` portfolio choice model was motivated by the finding that with high frequency or moderately high frequency data, means are more difficult to estimate than variances. -A model of **robust portfolio choice** that we'll describe below also begins -from the same starting point. +A model of **robust portfolio choice** that we'll describe below also begins from the same starting point. -To begin, we'll take for granted that means are more difficult to -estimate that covariances and will focus on how Black and Litterman, on -the one hand, an robust control theorists, on the other, would recommend -modifying the **mean-variance portfolio choice model** to take that into -account. +To begin, we'll take for granted that means are more difficult to estimate than covariances and will focus on how Black and Litterman, on the one hand, and robust control theorists, on the other, would recommend modifying the **mean-variance portfolio choice model** to take that into account. -At the end of this lecture, we shall use some rates of convergence -results and some simulations to verify how means are more difficult to -estimate than variances. +At the end of this lecture, we shall use some rates of convergence results and some simulations to verify how means are more difficult to estimate than variances. Among the ideas in play in this lecture will be @@ -63,20 +54,14 @@ Among the ideas in play in this lecture will be theory -In summary, we'll describe two ways to modify the classic -mean-variance portfolio choice model in ways designed to make its -recommendations more plausible. +In summary, we'll describe two ways to modify the classic mean-variance portfolio choice model in ways designed to make its recommendations more plausible. -Both of the adjustments that we describe are designed to confront a -widely recognized embarrassment to mean-variance portfolio theory, -namely, that it usually implies taking very extreme long-short portfolio -positions. +Both of the adjustments that we describe are designed to confront a widely recognized embarrassment to mean-variance portfolio theory, namely, that it usually implies taking very extreme long-short portfolio positions. -The two approaches build on a common and widespread hunch -- -that because it is much easier statistically to estimate covariances of -excess returns than it is to estimate their means, it makes sense to -adjust investors' subjective beliefs about mean returns in order to render more plausible decisions. +They tame those positions in different ways: the Black-Litterman adjustment can reverse which assets are held short, while the robust adjustment scales all positions toward zero without changing their signs. + +The two approaches build on a common and widespread hunch -- that because it is much easier statistically to estimate covariances of excess returns than it is to estimate their means, it makes sense to adjust investors' subjective beliefs about mean returns in order to render more plausible decisions. Let's start with some imports: @@ -92,13 +77,9 @@ from numba import jit A risk-free security earns one-period net return $r_f$. -An $n \times 1$ vector of risky securities earns -an $n \times 1$ vector $\vec r - r_f {\bf 1}$ of *excess -returns*, where ${\bf 1}$ is an $n \times 1$ vector of -ones. +An $n \times 1$ vector of risky securities earns an $n \times 1$ vector $\vec r - r_f {\bf 1}$ of *excess returns*, where ${\bf 1}$ is an $n \times 1$ vector of ones. -The excess return vector is multivariate normal with mean $\mu$ -and covariance matrix $\Sigma$, which we express either as +The excess return vector is multivariate normal with mean $\mu$ and covariance matrix $\Sigma$, which we express either as $$ \vec r - r_f {\bf 1} \sim {\mathcal N}(\mu, \Sigma) @@ -110,19 +91,17 @@ $$ \vec r - r_f {\bf 1} = \mu + C \epsilon $$ -where $\epsilon \sim {\mathcal N}(0, I)$ is an $n \times 1$ -random vector. +where $\epsilon \sim {\mathcal N}(0, I)$ is an $n \times 1$ random vector. -Let $w$ be an $n \times 1$ vector of portfolio weights. +Let $w$ be an $n \times 1$ vector of portfolio weights. -A portfolio consisting $w$ earns returns +A portfolio consisting of $w$ earns returns $$ w' (\vec r - r_f {\bf 1}) \sim {\mathcal N}(w' \mu, w' \Sigma w ) $$ -The **mean-variance portfolio choice problem** is to choose $w$ to -maximize +The **mean-variance portfolio choice problem** is to choose $w$ to maximize ```{math} :label: choice-problem @@ -130,8 +109,9 @@ maximize U(\mu,\Sigma;w) = w'\mu - \frac{\delta}{2} w' \Sigma w ``` -where $\delta > 0$ is a risk-aversion parameter. The first-order -condition for maximizing {eq}`choice-problem` with respect to the vector $w$ is +where $\delta > 0$ is a risk-aversion parameter. + +The first-order condition for maximizing {eq}`choice-problem` with respect to the vector $w$ is $$ \mu = \delta \Sigma w @@ -150,24 +130,16 @@ w = (\delta \Sigma)^{-1} \mu The key inputs into the portfolio choice model {eq}`risky-portfolio` are - estimates of the parameters $\mu, \Sigma$ of the random excess - return vector$(\vec r - r_f {\bf 1})$ + return vector $\vec r - r_f {\bf 1}$ - the risk-aversion parameter $\delta$ -A standard way of estimating $\mu$ is maximum-likelihood or least -squares; that amounts to estimating $\mu$ by a sample mean of -excess returns and estimating $\Sigma$ by a sample covariance -matrix. +A standard way of estimating $\mu$ is maximum-likelihood or least squares; that amounts to estimating $\mu$ by a sample mean of excess returns and estimating $\Sigma$ by a sample covariance matrix. ## Black-Litterman starting point -When estimates of $\mu$ and $\Sigma$ from historical -sample means and covariances have been combined with **plausible** values -of the risk-aversion parameter $\delta$ to compute an -optimal portfolio from formula {eq}`risky-portfolio`, a typical outcome has been -$w$'s with **extreme long and short positions**. +When estimates of $\mu$ and $\Sigma$ from historical sample means and covariances have been combined with **plausible** values of the risk-aversion parameter $\delta$ to compute an optimal portfolio from formula {eq}`risky-portfolio`, a typical outcome has been $w$'s with **extreme long and short positions**. -A common reaction to these outcomes is that they are so implausible that a portfolio -manager cannot recommend them to a customer. +A common reaction to these outcomes is that they are so implausible that a portfolio manager cannot recommend them to a customer. ```{code-cell} ipython3 np.random.seed(12) @@ -215,7 +187,7 @@ plt.legend(numpoints=1, fontsize=11) plt.show() ``` -Black and Litterman's responded to this situation in the following way: +Black and Litterman responded to this situation in the following way: - They continue to accept {eq}`risky-portfolio` as a good model for choosing an optimal portfolio $w$. @@ -226,11 +198,7 @@ Black and Litterman's responded to this situation in the following way: portfolio choices that are more plausible in terms of conforming to what most people actually do. -In particular, given $\Sigma$ and a plausible value -of $\delta$, Black and Litterman reverse engineered a vector -$\mu_{BL}$ of mean excess returns that makes the $w$ -implied by formula {eq}`risky-portfolio` equal the **actual** market portfolio -$w_m$, so that +In particular, given $\Sigma$ and a plausible value of $\delta$, Black and Litterman reverse engineered a vector $\mu_{BL}$ of mean excess returns that makes the $w$ implied by formula {eq}`risky-portfolio` equal the **actual** market portfolio $w_m$, so that $$ w_m = (\delta \Sigma)^{-1} \mu_{BL} @@ -252,8 +220,7 @@ $$ \sigma^2 = w_m' \Sigma w_m $$ -as the variance of the excess return on the market portfolio -$w_m$. +as the variance of the excess return on the market portfolio $w_m$. Define @@ -261,14 +228,11 @@ $$ {\bf SR}_m = \frac{ r_m - r_f}{\sigma} $$ -as the **Sharpe-ratio** on the market portfolio $w_m$. +as the **Sharpe ratio** of the market portfolio $w_m$. -Let $\delta_m$ be the value of the risk aversion parameter that -induces an investor to hold the market portfolio in light of the optimal -portfolio choice rule {eq}`risky-portfolio`. +Let $\delta_m$ be the value of the risk aversion parameter that induces an investor to hold the market portfolio in light of the optimal portfolio choice rule {eq}`risky-portfolio`. -Evidently, portfolio rule {eq}`risky-portfolio` then implies that -$r_m - r_f = \delta_m \sigma^2$ or +Evidently, portfolio rule {eq}`risky-portfolio` then implies that $r_m - r_f = \delta_m \sigma^2$ or $$ \delta_m = \frac{r_m - r_f}{\sigma^2} @@ -280,38 +244,32 @@ $$ \delta_m = \frac{{\bf SR}_m}{\sigma} $$ -Following the Black-Litterman philosophy, our first step will be to back -a value of $\delta_m$ from +Following the Black-Litterman philosophy, our first step will be to back out a value of $\delta_m$ from -- an estimate of the Sharpe-ratio, and +- an estimate of the Sharpe ratio, and - our maximum likelihood estimate of $\sigma$ drawn from our - estimates or $w_m$ and $\Sigma$ + estimates of $w_m$ and $\Sigma$ -The second key Black-Litterman step is then to use this value of -$\delta$ together with the maximum likelihood estimate of -$\Sigma$ to deduce a $\mu_{\bf BL}$ that verifies -portfolio rule {eq}`risky-portfolio` at the market portfolio $w = w_m$ +The second key Black-Litterman step is then to use this value of $\delta_m$ together with the maximum likelihood estimate of $\Sigma$ to deduce a $\mu_{BL}$ that verifies portfolio rule {eq}`risky-portfolio` at the market portfolio $w = w_m$ $$ -\mu_m = \delta_m \Sigma w_m +\mu_{BL} = \delta_m \Sigma w_m $$ -The starting point of the Black-Litterman portfolio choice model is thus -a pair $(\delta_m, \mu_m)$ that tells the customer to hold the -market portfolio. +The starting point of the Black-Litterman portfolio choice model is thus a pair $(\delta_m, \mu_{BL})$ that tells the customer to hold the market portfolio. ```{code-cell} ipython3 # Observed mean excess market return r_m = w_m @ μ_est # Estimated variance of the market portfolio -σ_m = w_m @ Σ_est @ w_m +var_m = w_m @ Σ_est @ w_m -# Sharpe-ratio -sr_m = r_m / np.sqrt(σ_m) +# Sharpe ratio +sr_m = r_m / np.sqrt(var_m) # Risk aversion of market portfolio holder -d_m = r_m / σ_m +d_m = r_m / var_m # Derive "view" which would induce the market portfolio μ_m = (d_m * Σ_est @ w_m).reshape(N, 1) @@ -331,9 +289,7 @@ plt.show() ## Adding views -Black and Litterman start with a baseline customer who asserts that he -or she shares the **market's views**, which means that he or she -believes that excess returns are governed by +Black and Litterman start with a baseline customer who asserts that he or she shares the **market's views**, which means that he or she believes that excess returns are governed by ```{math} :label: excess-returns @@ -341,28 +297,21 @@ believes that excess returns are governed by \vec r - r_f {\bf 1} \sim {\mathcal N}( \mu_{BL}, \Sigma) ``` -Black and Litterman would advise that customer to hold the market -portfolio of risky securities. +Black and Litterman would advise that customer to hold the market portfolio of risky securities. -Black and Litterman then imagine a consumer who would like to express a -view that differs from the market's. +Black and Litterman then imagine a customer who would like to express a view that differs from the market's. -The consumer wants appropriately to -mix his view with the market's before using {eq}`risky-portfolio` to choose a portfolio. +The customer wants appropriately to mix his view with the market's before using {eq}`risky-portfolio` to choose a portfolio. -Suppose that the customer's view is expressed by a hunch that rather -than {eq}`excess-returns`, excess returns are governed by +Suppose that the customer's view is expressed by a hunch that rather than {eq}`excess-returns`, excess returns are governed by $$ \vec r - r_f {\bf 1} \sim {\mathcal N}( \hat \mu, \tau \Sigma) $$ -where $\tau > 0$ is a scalar parameter that determines how the -decision maker wants to mix his view $\hat \mu$ with the market's -view $\mu_{\bf BL}$. +where $\tau > 0$ is a scalar parameter that determines how the decision maker wants to mix his view $\hat \mu$ with the market's view $\mu_{\bf BL}$. -Black and Litterman would then use a formula like the following one to -mix the views $\hat \mu$ and $\mu_{\bf BL}$ +Black and Litterman would then use a formula like the following one to mix the views $\hat \mu$ and $\mu_{\bf BL}$ ```{math} :label: mix-views @@ -370,26 +319,24 @@ mix the views $\hat \mu$ and $\mu_{\bf BL}$ \tilde \mu = (\Sigma^{-1} + (\tau \Sigma)^{-1})^{-1} (\Sigma^{-1} \mu_{BL} + (\tau \Sigma)^{-1} \hat \mu) ``` -Black and Litterman would then advise the customer to hold the portfolio -associated with these views implied by rule {eq}`risky-portfolio`: +Black and Litterman would then advise the customer to hold the portfolio associated with these views implied by rule {eq}`risky-portfolio`: $$ \tilde w = (\delta \Sigma)^{-1} \tilde \mu $$ -This portfolio $\tilde w$ will deviate from the -portfolio $w_{BL}$ in amounts that depend on the mixing parameter -$\tau$. +This portfolio $\tilde w$ will deviate from the market portfolio $w_m$ in amounts that depend on the mixing parameter $\tau$. -If $\hat \mu$ is the maximum likelihood estimator -and $\tau$ is chosen heavily to weight this view, then the -customer's portfolio will involve big short-long positions. +If $\hat \mu$ is the maximum likelihood estimator and $\tau$ is chosen heavily to weight this view, then the customer's portfolio will involve big short-long positions. ```{code-cell} ipython3 def black_litterman(λ, μ1, μ2, Σ1, Σ2): """ - This function calculates the Black-Litterman mixture - mean excess return and covariance matrix + Return the precision-weighted mixture of the means μ1 and μ2, + + (Σ1^{-1} + λ Σ2^{-1})^{-1} (Σ1^{-1} μ1 + λ Σ2^{-1} μ2), + + where λ scales the precision Σ2^{-1} of the second view. """ Σ1_inv = np.linalg.inv(Σ1) Σ2_inv = np.linalg.inv(Σ2) @@ -402,13 +349,14 @@ def black_litterman(λ, μ1, μ2, Σ1, Σ2): μ_tilde = black_litterman(1, μ_m, μ_est, Σ_est, τ * Σ_est) # The Black-Litterman recommendation for the portfolio weights -w_tilde = np.linalg.solve(δ * Σ_est, μ_tilde) +# (use the market-implied risk aversion d_m, so that μ_m reproduces w_m) +w_tilde = np.linalg.solve(d_m * Σ_est, μ_tilde) ``` ```{code-cell} ipython3 def BL_plot(τ): μ_tilde = black_litterman(1, μ_m, μ_est, Σ_est, τ * Σ_est) - w_tilde = np.linalg.solve(δ * Σ_est, μ_tilde) + w_tilde = np.linalg.solve(d_m * Σ_est, μ_tilde) fig, ax = plt.subplots(1, 2, figsize=(16, 6)) ax[0].plot(np.arange(N)+1, μ_est, 'o', c='k', @@ -420,7 +368,7 @@ def BL_plot(τ): ax[0].vlines(np.arange(N)+1, μ_m, μ_est, lw=1) ax[0].axhline(0, c='k', ls='--') ax[0].set(xlim=(0, N+1), xlabel='Assets', - title=r'Relationship between $\hat{\mu}$, $\mu_{BL}$, and $ \tilde{\mu}$') + title=r'Relationship between $\hat{\mu}$, $\mu_{BL}$, and $\tilde{\mu}$') ax[0].xaxis.set_ticks(np.arange(1, N+1, 1)) ax[0].legend(numpoints=1) @@ -445,33 +393,31 @@ BL_plot(τ) ## Bayesian interpretation -Consider the following Bayesian interpretation of the Black-Litterman -recommendation. +Consider the following Bayesian interpretation of the Black-Litterman recommendation. -The prior belief over the mean excess returns is consistent with the -market portfolio and is given by +The prior belief over the mean excess returns is consistent with the market portfolio and is given by $$ \mu \sim \mathcal{N}(\mu_{BL}, \Sigma) $$ -Given a particular realization of the mean excess returns -$\mu$ one observes the average excess returns $\hat \mu$ -on the market according to the distribution +Given a particular realization of the mean excess returns $\mu$ one observes the average excess returns $\hat \mu$ on the market according to the distribution $$ \hat \mu \mid \mu, \Sigma \sim \mathcal{N}(\mu, \tau\Sigma) $$ -where $\tau$ is typically small capturing the idea that the -variation in the mean is smaller than the variation of the individual -random variable. +where $\tau$ scales the uncertainty of the investor's own estimate $\hat \mu$ relative to the prior. + +If $\hat \mu$ is the sample mean of $T$ i.i.d. observations, the natural value is $\tau = 1/T$, which puts almost all weight on $\hat \mu$ and so leads back to extreme long-short positions. + +That is why practitioners following He and Litterman instead attach a small scalar to the prior, $\mu \sim \mathcal{N}(\mu_{BL}, \tau \Sigma)$, thereby expressing confidence in the market's view. -Given the realized excess returns one should then update the prior over -the mean excess returns according to Bayes rule. +Exercise {ref}`bl_ex1` compares the two conventions. -The corresponding -posterior over mean excess returns is normally distributed with mean +Given the realized excess returns one should then update the prior over the mean excess returns according to Bayes' rule. + +The corresponding posterior over mean excess returns is normally distributed with mean $$ (\Sigma^{-1} + (\tau \Sigma)^{-1})^{-1} (\Sigma^{-1}\mu_{BL} + (\tau \Sigma)^{-1} \hat \mu) @@ -483,9 +429,7 @@ $$ (\Sigma^{-1} + (\tau \Sigma)^{-1})^{-1} $$ -Hence, the Black-Litterman recommendation is consistent with the Bayes -update of the prior over the mean excess returns in light of the -realized average excess returns on the market. +Hence, the Black-Litterman recommendation is consistent with the Bayes update of the prior over the mean excess returns in light of the realized average excess returns on the market. ## Curve decolletage @@ -501,12 +445,9 @@ $$ \vec r_e \sim {\mathcal N}( \hat{\mu}, \tau\Sigma) $$ -A special feature of the multivariate normal random variable -$Z$ is that its density function depends only on the (Euclidiean) -length of its realization $z$. +A special feature of the multivariate normal random variable $Z$ is that, after standardization, its density function depends only on the (Euclidean) length of the standardized realization. -Formally, let the -$k$-dimensional random vector be +Formally, let the $k$-dimensional random vector be $$ Z\sim \mathcal{N}(\mu, \Sigma) @@ -515,11 +456,10 @@ $$ then $$ -\bar{Z} \equiv \Sigma(Z-\mu)\sim \mathcal{N}(\mathbf{0}, I) +\bar{Z} \equiv \Sigma^{-1/2}(Z-\mu)\sim \mathcal{N}(\mathbf{0}, I) $$ -and so the points where the density takes the same value can be -described by the ellipse +and so the points where the density takes the same value can be described by the ellipse (an ellipsoid when $k > 2$) ```{math} :label: ellipse @@ -527,11 +467,9 @@ described by the ellipse \bar z \cdot \bar z = (z - \mu)'\Sigma^{-1}(z - \mu) = \bar d ``` -where $\bar d\in\mathbb{R}_+$ denotes the (transformation) of a -particular density value. +where $\bar d\in\mathbb{R}_+$ denotes the (transformation) of a particular density value. -The curves defined by equation {eq}`ellipse` can be -labeled as iso-likelihood ellipses +The curves defined by equation {eq}`ellipse` can be labeled as iso-likelihood ellipses > **Remark:** More generally there is a class of density functions > that possesses this feature, i.e. @@ -542,13 +480,10 @@ labeled as iso-likelihood ellipses \text{ has the form } \quad f(z) = c g(z\cdot z) $$ > -> This property is called **spherical symmetry** (see p 81. in Leamer +> This property is called **spherical symmetry** (see p. 81 of Leamer > (1978) {cite}`leamer1978specification`). -In our specific example, we can use the pair -$(\bar d_1, \bar d_2)$ as being two "likelihood" values for which -the corresponding iso-likelihood ellipses in the excess return space are -given by +In our specific example, we can use the pair $(\bar d_1, \bar d_2)$ as being two "likelihood" values for which the corresponding iso-likelihood ellipses in the excess return space are given by $$ \begin{aligned} @@ -557,43 +492,34 @@ $$ \end{aligned} $$ -Notice that for particular $\bar d_1$ and $\bar d_2$ values -the two ellipses have a tangency point. +Notice that for particular $\bar d_1$ and $\bar d_2$ values the two ellipses have a tangency point. + +These tangency points, indexed by the pairs $(\bar d_1, \bar d_2)$, characterize points $\vec r_e$ from which there exists no deviation where one can increase the likelihood of one view without decreasing the likelihood of the other view. -These tangency points, indexed -by the pairs $(\bar d_1, \bar d_2)$, characterize points -$\vec r_e$ from which there exists no deviation where one can -increase the likelihood of one view without decreasing the likelihood of -the other view. +The pairs $(\bar d_1, \bar d_2)$ for which there is such a point outline a curve in the excess return space. -The pairs $(\bar d_1, \bar d_2)$ for which there -is such a point outlines a curve in the excess return space. This curve -is reminiscent of the Pareto curve in an Edgeworth-box setting. +This curve is reminiscent of the Pareto curve in an Edgeworth-box setting. Dickey (1975) {cite}`Dickey1975` calls it a *curve decolletage*. -Leamer (1978) {cite}`leamer1978specification` calls it an *information contract curve* and -describes it by the following program: maximize the likelihood of one -view, say the Black-Litterman recommendation while keeping the -likelihood of the other view at least at a prespecified constant -$\bar d_2$ +Leamer (1978) {cite}`leamer1978specification` calls it an *information contract curve* and describes it by the following program: maximize the likelihood of one view, say the Black-Litterman recommendation, while keeping the likelihood of the other view at least at a prespecified level indexed by $\bar d_2$. + +Because each quadratic form below is, up to constants, minus twice the log likelihood of the corresponding view, this amounts to minimizing the first quadratic form subject to the second quadratic form being at most $\bar d_2$ $$ \begin{aligned} - \bar d_1(\bar d_2) &\equiv \max_{\vec r_e} \ \ (\vec r_e - \mu_{BL})'\Sigma^{-1}(\vec r_e - \mu_{BL}) \\ -\text{subject to } \quad &(\vec r_e - \hat\mu)'(\tau\Sigma)^{-1}(\vec r_e - \hat \mu) \geq \bar d_2 + \bar d_1(\bar d_2) &\equiv \min_{\vec r_e} \ \ (\vec r_e - \mu_{BL})'\Sigma^{-1}(\vec r_e - \mu_{BL}) \\ +\text{subject to } \quad &(\vec r_e - \hat\mu)'(\tau\Sigma)^{-1}(\vec r_e - \hat \mu) \leq \bar d_2 \end{aligned} $$ -Denoting the multiplier on the constraint by $\lambda$, the -first-order condition is +Denoting the multiplier on the constraint by $\lambda \geq 0$, the first-order condition is $$ 2(\vec r_e - \mu_{BL} )'\Sigma^{-1} + \lambda 2(\vec r_e - \hat\mu)'(\tau\Sigma)^{-1} = \mathbf{0} $$ -which defines the *information contract curve* between -$\mu_{BL}$ and $\hat \mu$ +which defines the *information contract curve* between $\mu_{BL}$ and $\hat \mu$ ```{math} :label: info-curve @@ -601,13 +527,9 @@ $\mu_{BL}$ and $\hat \mu$ \vec r_e = (\Sigma^{-1} + \lambda (\tau \Sigma)^{-1})^{-1} (\Sigma^{-1} \mu_{BL} + \lambda (\tau \Sigma)^{-1}\hat \mu ) ``` -Note that if $\lambda = 1$, {eq}`info-curve` is equivalent with {eq}`mix-views` and it -identifies one point on the information contract curve. +Note that if $\lambda = 1$, {eq}`info-curve` is equivalent to {eq}`mix-views` and it identifies one point on the information contract curve. -Furthermore, because $\lambda$ is a function of the minimum likelihood -$\bar d_2$ on the RHS of the constraint, by varying -$\bar d_2$ (or $\lambda$ ), we can trace out the whole curve -as the figure below illustrates. +Furthermore, because $\lambda$ is a function of the minimum likelihood $\bar d_2$ on the RHS of the constraint, by varying $\bar d_2$ (or $\lambda$ ), we can trace out the whole curve as the figure below illustrates. ```{code-cell} ipython3 np.random.seed(1987102) @@ -631,8 +553,8 @@ sample = excess_return.rvs(T) μ_est = sample.mean(0).reshape(N, 1) Σ_est = np.cov(sample.T) -σ_m = w_m @ Σ_est @ w_m -d_m = (w_m @ μ_est) / σ_m +var_m = w_m @ Σ_est @ w_m +d_m = (w_m @ μ_est) / var_m μ_m = (d_m * Σ_est @ w_m).reshape(N, 1) N_r1, N_r2 = 100, 100 @@ -653,7 +575,7 @@ def decolletage(λ): X, Y = np.meshgrid(r1, r2) XY = np.stack((X, Y), axis=-1) - Z_BL = dist_r_BL.pdf(XY) + Z_BL = dist_r_BL.pdf(XY) Z_hat = dist_r_hat.pdf(XY) μ_tilde = black_litterman(λ, μ_m, μ_est, Σ_est, τ * Σ_est).flatten() @@ -679,19 +601,11 @@ def decolletage(λ): decolletage(λ) ``` -Note that the line that connects the two points -$\hat \mu$ and $\mu_{BL}$ is linear, which comes from the -fact that the covariance matrices of the two competing distributions -(views) are proportional to each other. +Note that the line that connects the two points $\hat \mu$ and $\mu_{BL}$ is linear, which comes from the fact that the covariance matrices of the two competing distributions (views) are proportional to each other. -To illustrate the fact that this is not necessarily the case, consider -another example using the same parameter values, except that the "second -view" constituting the constraint has covariance matrix -$\tau I$ instead of $\tau \Sigma$. +To illustrate the fact that this is not necessarily the case, consider another example using the same parameter values, except that the "second view" constituting the constraint has covariance matrix $\tau I$ instead of $\tau \Sigma$. -This leads to the -following figure, on which the curve connecting $\hat \mu$ -and $\mu_{BL}$ are bending +This leads to the following figure, on which the curve connecting $\hat \mu$ and $\mu_{BL}$ bends ```{code-cell} ipython3 λ_grid = np.linspace(.001, 20000, 1000) @@ -707,7 +621,7 @@ def decolletage(λ): X, Y = np.meshgrid(r1, r2) XY = np.stack((X, Y), axis=-1) - Z_BL = dist_r_BL.pdf(XY) + Z_BL = dist_r_BL.pdf(XY) Z_hat = dist_r_hat.pdf(XY) μ_tilde = black_litterman(λ, μ_m, μ_est, Σ_est, τ * np.eye(N)).flatten() @@ -748,46 +662,36 @@ $$ \hat{\beta}_{OLS} = (X'X)^{-1}X'y $$ -A common performance measure of estimators is the *mean squared error -(MSE)*. +A common performance measure of estimators is the *mean squared error (MSE)*. -An estimator is "good" if its MSE is relatively small. Suppose -that $\beta_0$ is the "true" value of the coefficient, then the MSE -of the OLS estimator is +An estimator is "good" if its MSE is relatively small. + +Suppose that $\beta_0$ is the "true" value of the coefficient, then the MSE of the OLS estimator is $$ \text{mse}(\hat \beta_{OLS}, \beta_0) := \mathbb E \Vert \hat \beta_{OLS} - \beta_0\Vert^2 = \underbrace{\mathbb E \Vert \hat \beta_{OLS} - \mathbb E -\beta_{OLS}\Vert^2}_{\text{variance}} + -\underbrace{\Vert \mathbb E \hat\beta_{OLS} - \beta_0\Vert^2}_{\text{bias}} +\hat \beta_{OLS}\Vert^2}_{\text{variance}} + +\underbrace{\Vert \mathbb E \hat\beta_{OLS} - \beta_0\Vert^2}_{\text{squared bias}} $$ -From this decomposition, one can see that in order for the MSE to be -small, both the bias and the variance terms must be small. +From this decomposition, one can see that in order for the MSE to be small, both the bias and the variance terms must be small. -For example, -consider the case when $X$ is a $T$-vector of ones (where -$T$ is the sample size), so $\hat\beta_{OLS}$ is simply the -sample average, while $\beta_0\in \mathbb{R}$ is defined by the -true mean of $y$. +For example, consider the case when $X$ is a $T$-vector of ones (where $T$ is the sample size), so $\hat\beta_{OLS}$ is simply the sample average, while $\beta_0\in \mathbb{R}$ is defined by the true mean of $y$. In this example the MSE is $$ \text{mse}(\hat \beta_{OLS}, \beta_0) = \underbrace{\frac{1}{T^2} \mathbb E \left(\sum_{t=1}^{T} (y_{t}- \beta_0)\right)^2 }_{\text{variance}} + -\underbrace{0}_{\text{bias}} +\underbrace{0}_{\text{squared bias}} $$ -However, because there is a trade-off between the estimator's bias and -variance, there are cases when by permitting a small bias we can -substantially reduce the variance so overall the MSE gets smaller. +However, because there is a trade-off between the estimator's bias and variance, there are cases when by permitting a small bias we can substantially reduce the variance so overall the MSE gets smaller. -A typical scenario when this proves to be useful is when the number of -coefficients to be estimated is large relative to the sample size. +A typical scenario when this proves to be useful is when the number of coefficients to be estimated is large relative to the sample size. -In these cases, one approach to handle the bias-variance trade-off is the -so called *Tikhonov regularization*. +In these cases, one approach to handle the bias-variance trade-off is the so called *Tikhonov regularization*. A general form with regularization matrix $\Gamma$ can be written as @@ -807,34 +711,27 @@ $$ \hat{\beta}_{Reg} = (X'X + \Gamma'\Gamma)^{-1}(X'X\hat{\beta}_{OLS} + \Gamma'\Gamma\tilde \beta) $$ -Often, the regularization matrix takes the form -$\Gamma = \lambda I$ with $\lambda>0$ -and $\tilde \beta = \mathbf{0}$. +Often, the regularization matrix takes the form $\Gamma = \sqrt{\lambda} I$ with $\lambda>0$ and $\tilde \beta = \mathbf{0}$. Then the Tikhonov regularization is equivalent to what is called *ridge regression* in statistics. -To illustrate how this estimator addresses the bias-variance trade-off, -we compute the MSE of the ridge estimator +To illustrate how this estimator addresses the bias-variance trade-off, we compute the MSE of the ridge estimator $$ \text{mse}(\hat \beta_{\text{ridge}}, \beta_0) = \underbrace{\frac{1}{(T+\lambda)^2} \mathbb E \left(\sum_{t=1}^{T} (y_{t}- \beta_0)\right)^2 }_{\text{variance}} + -\underbrace{\left(\frac{\lambda}{T+\lambda}\right)^2 \beta_0^2}_{\text{bias}} +\underbrace{\left(\frac{\lambda}{T+\lambda}\right)^2 \beta_0^2}_{\text{squared bias}} $$ -The ridge regression shrinks the coefficients of the estimated vector -towards zero relative to the OLS estimates thus reducing the variance -term at the cost of introducing a "small" bias. +The ridge regression shrinks the coefficients of the estimated vector towards zero relative to the OLS estimates thus reducing the variance term at the cost of introducing a "small" bias. However, there is nothing special about the zero vector. -When $\tilde \beta \neq \mathbf{0}$ shrinkage occurs in the direction -of $\tilde \beta$. +When $\tilde \beta \neq \mathbf{0}$ shrinkage occurs in the direction of $\tilde \beta$. -Now, we can give a regularization interpretation of the Black-Litterman -portfolio recommendation. +Now, we can give a regularization interpretation of the Black-Litterman portfolio recommendation. -To this end, first simplify the equation {eq}`mix-views` that characterizes the Black-Litterman recommendation +To this end, first simplify the equation {eq}`mix-views` that characterizes the Black-Litterman recommendation $$ \begin{aligned} @@ -844,67 +741,51 @@ $$ \end{aligned} $$ -In our case, $\hat \mu$ is the estimated mean excess returns of -securities. This could be written as a vector autoregression where +In our case, $\hat \mu$ is the vector of estimated mean excess returns of securities. + +It can be computed from a stacked linear regression of returns on constants (a system of seemingly unrelated regressions with the same regressors in every equation) where - $y$ is the stacked vector of observed excess returns of size $(N T\times 1)$ -- $N$ securities and $T$ observations. -- $X = \sqrt{T^{-1}}(I_{N} \otimes \iota_T)$ where $I_N$ +- $X = I_{N} \otimes \iota_T$ where $I_N$ is the identity matrix and $\iota_T$ is a column vector of ones. -Correspondingly, the OLS regression of $y$ on $X$ would -yield the mean excess returns as coefficients. +Correspondingly, the OLS regression of $y$ on $X$ yields $\hat \beta_{OLS} = (X'X)^{-1} X'y = \hat \mu$, the vector of mean excess returns. -With $\Gamma = \sqrt{\tau T^{-1}}(I_{N} \otimes \iota_T)$ we can -write the regularized version of the mean excess return estimation +With $\Gamma = \sqrt{\tau}(I_{N} \otimes \iota_T)$, so that $\Gamma'\Gamma = \tau X'X = \tau T I_N$, we can write the regularized version of the mean excess return estimation $$ \begin{aligned} \hat{\beta}_{Reg} &= (X'X + \Gamma'\Gamma)^{-1}(X'X\hat{\beta}_{OLS} + \Gamma'\Gamma\tilde \beta) \\ -&= (1 + \tau)^{-1}X'X (X'X)^{-1} (\hat \beta_{OLS} + \tau \tilde \beta) \\ +&= \left((1 + \tau) T I_N\right)^{-1} T (\hat \beta_{OLS} + \tau \tilde \beta) \\ &= (1 + \tau)^{-1} (\hat \beta_{OLS} + \tau \tilde \beta) \\ &= (1 + \tau^{-1})^{-1} ( \tau^{-1}\hat \beta_{OLS} + \tilde \beta) \end{aligned} $$ -Given that -$\hat \beta_{OLS} = \hat \mu$ and $\tilde \beta = \mu_{BL}$ -in the Black-Litterman model, we have the following interpretation of the -model's recommendation. +Given that $\hat \beta_{OLS} = \hat \mu$ and $\tilde \beta = \mu_{BL}$ in the Black-Litterman model, we have the following interpretation of the model's recommendation. -The estimated (personal) view of the mean excess returns, -$\hat{\mu}$ that would lead to extreme short-long positions are -"shrunk" towards the conservative market view, $\mu_{BL}$, that -leads to the more conservative market portfolio. +The estimated (personal) view of the mean excess returns, $\hat{\mu}$ that would lead to extreme short-long positions are "shrunk" towards the conservative market view, $\mu_{BL}$, that leads to the more conservative market portfolio. -So the Black-Litterman procedure results in a recommendation that is a -compromise between the conservative market portfolio and the more -extreme portfolio that is implied by estimated "personal" views. +So the Black-Litterman procedure results in a recommendation that is a compromise between the conservative market portfolio and the more extreme portfolio that is implied by estimated "personal" views. ## A robust control operator -The Black-Litterman approach is partly inspired by the econometric -insight that it is easier to estimate covariances of excess returns than -the means. +The Black-Litterman approach is partly inspired by the econometric insight that it is easier to estimate covariances of excess returns than the means. -That is what gave Black and Litterman license to adjust -investors' perception of mean excess returns while not tampering with -the covariance matrix of excess returns. +That is what gave Black and Litterman license to adjust investors' perception of mean excess returns while not tampering with the covariance matrix of excess returns. -The robust control theory is another approach that also hinges on -adjusting mean excess returns but not covariances. +The robust control theory is another approach that also hinges on adjusting mean excess returns but not covariances. -Associated with a robust control problem is what Hansen and Sargent {cite}`HansenSargent2001`, {cite}`HansenSargent2008` call -a ${\sf T}$ operator. +Associated with a robust control problem is what Hansen and Sargent {cite}`HansenSargent2001`, {cite}`HansenSargent2008` call a ${\sf T}$ operator. -Let's define the ${\sf T}$ operator as it applies to the problem -at hand. +Let's define the ${\sf T}$ operator as it applies to the problem at hand. -Let $x$ be an $n \times 1$ Gaussian random vector with mean -vector $\mu$ and covariance matrix $\Sigma = C C'$. This -means that $x$ can be represented as +Let $x$ be an $n \times 1$ Gaussian random vector with mean vector $\mu$ and covariance matrix $\Sigma = C C'$. + +This means that $x$ can be represented as $$ x = \mu + C \epsilon @@ -912,30 +793,24 @@ $$ where $\epsilon \sim {\mathcal N}(0,I)$. -Let $\phi(\epsilon)$ denote the associated standardized Gaussian -density. +Let $\phi(\epsilon)$ denote the associated standardized Gaussian density. -Let $m(\epsilon,\mu)$ be a **likelihood ratio**, meaning that it -satisfies +Let $m(\epsilon,\mu)$ be a **likelihood ratio**, meaning that it satisfies - $m(\epsilon, \mu) > 0$ - $\int m(\epsilon,\mu) \phi(\epsilon) d \epsilon =1$ -That is, $m(\epsilon, \mu)$ is a non-negative random variable with -mean 1. +That is, $m(\epsilon, \mu)$ is a non-negative random variable with mean 1. -Multiplying $\phi(\epsilon)$ -by the likelihood ratio $m(\epsilon, \mu)$ produces a distorted distribution for -$\epsilon$, namely +Multiplying $\phi(\epsilon)$ by the likelihood ratio $m(\epsilon, \mu)$ produces a distorted distribution for $\epsilon$, namely $$ \tilde \phi(\epsilon) = m(\epsilon,\mu) \phi(\epsilon) $$ -The next concept that we need is the **entropy** of the distorted -distribution $\tilde \phi$ with respect to $\phi$. +The next concept that we need is the **relative entropy** of the distorted distribution $\tilde \phi$ with respect to $\phi$. -**Entropy** is defined as +**Relative entropy** is defined as $$ {\rm ent} = \int \log m(\epsilon,\mu) m(\epsilon,\mu) \phi(\epsilon) d \epsilon @@ -947,16 +822,13 @@ $$ {\rm ent} = \int \log m(\epsilon,\mu) \tilde \phi(\epsilon) d \epsilon $$ -That is, relative entropy is the expected value of the likelihood ratio -$m$ where the expectation is taken with respect to the twisted -density $\tilde \phi$. +That is, relative entropy is the expected value of the log likelihood ratio $\log m$, where the expectation is taken with respect to the twisted density $\tilde \phi$. + +Relative entropy is non-negative. -Relative entropy is non-negative. It is a measure of the discrepancy -between two probability distributions. +It is a measure of the discrepancy between two probability distributions. -As such, it plays an important -role in governing the behavior of statistical tests designed to -discriminate one probability distribution from another. +As such, it plays an important role in governing the behavior of statistical tests designed to discriminate one probability distribution from another. We are ready to define the ${\sf T}$ operator. @@ -966,14 +838,10 @@ Define $$ \begin{aligned} {\sf T}\left(V(x)\right) & = \min_{m(\epsilon,\mu)} \int m(\epsilon,\mu)[V(\mu + C \epsilon) + \theta \log m(\epsilon,\mu) ] \phi(\epsilon) d \epsilon \cr - & = - \log \theta \int \exp \left( \frac{- V(\mu + C \epsilon)}{\theta} \right) \phi(\epsilon) d \epsilon \end{aligned} + & = - \theta \log \int \exp \left( \frac{- V(\mu + C \epsilon)}{\theta} \right) \phi(\epsilon) d \epsilon \end{aligned} $$ -This asserts that ${\sf T}$ is an indirect utility function for a -minimization problem in which an **adversary** chooses a distorted -probability distribution $\tilde \phi$ to lower expected utility, -subject to a penalty term that gets bigger the larger is relative -entropy. +This asserts that ${\sf T}$ is an indirect utility function for a minimization problem in which an **adversary** chooses a distorted probability distribution $\tilde \phi$ to lower expected utility, subject to a penalty term that gets bigger the larger is relative entropy. Here the penalty parameter @@ -981,8 +849,9 @@ $$ \theta \in [\underline \theta, +\infty] $$ -is a robustness parameter when it is $+\infty$, there is no scope for the minimizing agent to distort the distribution, -so no robustness to alternative distributions is acquired. +is a robustness parameter. + +When $\theta = +\infty$, there is no scope for the minimizing agent to distort the distribution, so no robustness to alternative distributions is acquired. As $\theta$ is lowered, more robustness is achieved. @@ -991,26 +860,21 @@ The ${\sf T}$ operator is sometimes called a *risk-sensitivity* operator. ``` -We shall apply ${\sf T}$ to the special case of a linear value -function $w'(\vec r - r_f 1)$ -where $\vec r - r_f 1 \sim {\mathcal N}(\mu,\Sigma)$ or -$\vec r - r_f {\bf 1} = \mu + C \epsilon$ and -$\epsilon \sim {\mathcal N}(0,I)$. +We shall apply ${\sf T}$ to the special case of a linear value function $w'(\vec r - r_f {\bf 1})$ where $\vec r - r_f {\bf 1} \sim {\mathcal N}(\mu,\Sigma)$ or $\vec r - r_f {\bf 1} = \mu + C \epsilon$ and $\epsilon \sim {\mathcal N}(0,I)$. -The associated worst-case distribution of $\epsilon$ is Gaussian -with mean $v =-\theta^{-1} C' w$ and covariance matrix $I$ +The associated worst-case distribution of $\epsilon$ is Gaussian with mean $v =-\theta^{-1} C' w$ and covariance matrix $I$. (When the value function is affine, the worst-case distribution distorts the mean vector of $\epsilon$ but not the covariance matrix -of $\epsilon$). +of $\epsilon$.) -For utility function argument $w'(\vec r - r_f 1)$ +For utility function argument $w'(\vec r - r_f {\bf 1})$ $$ -{\sf T} ( \vec r - r_f {\bf 1}) = w' \mu + \zeta - \frac{1}{2 \theta} w' \Sigma w +{\sf T} \left( w'(\vec r - r_f {\bf 1}) \right) = w' \mu - \frac{1}{2 \theta} w' \Sigma w $$ -and entropy is +and relative entropy is $$ \frac{v'v}{2} = \frac{1}{2\theta^2} w' C C' w @@ -1018,11 +882,10 @@ $$ ## A robust mean-variance portfolio model -According to criterion {eq}`choice-problem`, the mean-variance portfolio choice problem -chooses $w$ to maximize +According to criterion {eq}`choice-problem`, the mean-variance portfolio choice problem chooses $w$ to maximize $$ -E [w ( \vec r - r_f {\bf 1})]] - {\rm var} [ w ( \vec r - r_f {\bf 1}) ] +E [w' ( \vec r - r_f {\bf 1})] - \frac{\delta}{2} {\rm var} [ w' ( \vec r - r_f {\bf 1}) ] $$ which equals @@ -1031,33 +894,32 @@ $$ w'\mu - \frac{\delta}{2} w' \Sigma w $$ -A robust decision maker can be modeled as replacing the mean return -$E [w ( \vec r - r_f {\bf 1})]$ with the risk-sensitive criterion +A robust decision maker can be modeled as replacing the mean return $E [w' ( \vec r - r_f {\bf 1})]$ with the risk-sensitive criterion $$ -{\sf T} [w ( \vec r - r_f {\bf 1})] = w' \mu - \frac{1}{2 \theta} w' \Sigma w +{\sf T} [w' ( \vec r - r_f {\bf 1})] = w' \mu - \frac{1}{2 \theta} w' \Sigma w $$ -that comes from replacing the mean $\mu$ of $\vec r - r\_f {\bf 1}$ with the worst-case mean +that comes from replacing the mean $\mu$ of $\vec r - r_f {\bf 1}$ with the worst-case mean $$ \mu - \theta^{-1} \Sigma w $$ -Notice how the worst-case mean vector depends on the portfolio -$w$. +and adding back the entropy penalty $\theta \frac{v'v}{2} = \frac{1}{2\theta} w' \Sigma w$, so that + +$$ +{\sf T} [w' ( \vec r - r_f {\bf 1})] = w' (\mu - \theta^{-1} \Sigma w ) + \frac{1}{2\theta} w' \Sigma w +$$ + +Notice how the worst-case mean vector depends on the portfolio $w$. -The operator ${\sf T}$ is the indirect utility function that -emerges from solving a problem in which an agent who chooses -probabilities does so in order to minimize the expected utility of a -maximizing agent (in our case, the maximizing agent chooses portfolio -weights $w$). +The operator ${\sf T}$ is the indirect utility function that emerges from solving a problem in which an agent who chooses probabilities does so in order to minimize the expected utility of a maximizing agent (in our case, the maximizing agent chooses portfolio weights $w$). -The robust version of the mean-variance portfolio choice problem is then -to choose a portfolio $w$ that maximizes +The robust version of the mean-variance portfolio choice problem is then to choose a portfolio $w$ that maximizes $$ -{\sf T} [w ( \vec r - r_f {\bf 1})] - \frac{\delta}{2} w' \Sigma w +{\sf T} [w' ( \vec r - r_f {\bf 1})] - \frac{\delta}{2} w' \Sigma w $$ or @@ -1065,47 +927,42 @@ or ```{math} :label: robust-mean-variance -w' (\mu - \theta^{-1} \Sigma w ) - \frac{\delta}{2} w' \Sigma w +w' \mu - \frac{\gamma}{2} w' \Sigma w - \frac{\delta}{2} w' \Sigma w ``` -The minimizer of {eq}`robust-mean-variance` is +The maximizer of {eq}`robust-mean-variance` is $$ w_{\rm rob} = \frac{1}{\delta + \gamma } \Sigma^{-1} \mu $$ -where $\gamma \equiv \theta^{-1}$ is sometimes called the -risk-sensitivity parameter. +where $\gamma \equiv \theta^{-1}$ is sometimes called the risk-sensitivity parameter. + +An increase in the risk-sensitivity parameter $\gamma$ shrinks the portfolio weights toward zero in the same way that an increase in risk aversion does. + +Indeed, $w_{\rm rob} = \frac{\delta}{\delta + \gamma} w$, a scalar multiple of the mean-variance portfolio $w = (\delta \Sigma)^{-1} \mu$ in {eq}`risky-portfolio`. + +Robustness therefore shrinks the sizes of long and short positions but leaves their signs unchanged: the same assets are shorted. -An increase in the risk-sensitivity parameter $\gamma$ shrinks the -portfolio weights toward zero in the same way that an increase in risk -aversion does. +In contrast, shrinking $\hat \mu$ toward $\mu_{BL}$, as Black and Litterman do, can reverse the sign of a position. ## Appendix -We want to illustrate the "folk theorem" that with high or moderate -frequency data, it is more difficult to estimate means than variances. +We want to illustrate the "folk theorem" that with high or moderate frequency data, it is more difficult to estimate means than variances. -In order to operationalize this statement, we take two analog -estimators: +In order to operationalize this statement, we take two analog estimators: - sample average: $\bar X_N = \frac{1}{N}\sum_{i=1}^{N} X_i$ - sample variance: - $S_N = \frac{1}{N-1}\sum_{t=1}^{N} (X_i - \bar X_N)^2$ + $S_N = \frac{1}{N-1}\sum_{i=1}^{N} (X_i - \bar X_N)^2$ -to estimate the unconditional mean and unconditional variance of the -random variable $X$, respectively. +to estimate the unconditional mean and unconditional variance of the random variable $X$, respectively. -To measure the "difficulty of estimation", we use *mean squared error* -(MSE), that is the average squared difference between the estimator and -the true value. +To measure the "difficulty of estimation", we use *mean squared error* (MSE), that is the average squared difference between the estimator and the true value. -Assuming that the process $\{X_i\}$is ergodic, -both analog estimators are known to converge to their true values as the -sample size $N$ goes to infinity. +Assuming that the process $\{X_i\}$ is ergodic, both analog estimators are known to converge to their true values as the sample size $N$ goes to infinity. -More precisely -for all $\varepsilon > 0$ +More precisely for all $\varepsilon > 0$ $$ \lim_{N\to \infty} \ \ P\left\{ \left |\bar X_N - \mathbb E X \right| > \varepsilon \right\} = 0 \quad \quad @@ -1117,16 +974,15 @@ $$ \lim_{N\to \infty} \ \ P \left\{ \left| S_N - \mathbb V X \right| > \varepsilon \right\} = 0 $$ -A necessary condition for these convergence results is that the -associated MSEs vanish as $N$ goes to infinity, or in other words, +A sufficient condition for these convergence results is that the associated MSEs vanish as $N$ goes to infinity, or in other words, $$ \text{MSE}(\bar X_N, \mathbb E X) = o(1) \quad \quad \text{and} \quad \quad \text{MSE}(S_N, \mathbb V X) = o(1) $$ -Even if the MSEs converge to zero, the associated rates might be -different. Looking at the limit of the *relative MSE* (as the sample -size grows to infinity) +Even if the MSEs converge to zero, the associated rates might be different. + +Looking at the limit of the *relative MSE* (as the sample size grows to infinity) $$ \frac{\text{MSE}(S_N, \mathbb V X)}{\text{MSE}(\bar X_N, \mathbb E X)} = \frac{o(1)}{o(1)} \underset{N \to \infty}{\to} B @@ -1134,28 +990,19 @@ $$ can inform us about the relative (asymptotic) rates. -We will show that in general, with dependent data, the limit -$B$ depends on the sampling frequency. +We will show that in general, with dependent data, the limit $B$ depends on the sampling frequency. -In particular, we find -that the rate of convergence of the variance estimator is less sensitive -to increased sampling frequency than the rate of convergence of the mean -estimator. +In particular, we find that the rate of convergence of the variance estimator is less sensitive to increased sampling frequency than the rate of convergence of the mean estimator. -Hence, we can expect the relative asymptotic -rate, $B$, to get smaller with higher frequency data, -illustrating that "it is more difficult to estimate means than -variances". +Hence, we can expect the relative asymptotic rate, $B$, to get smaller with higher frequency data, illustrating that "it is more difficult to estimate means than variances". -That is, we need significantly more data to obtain a given -precision of the mean estimate than for our variance estimate. +That is, we need significantly more data to obtain a given precision of the mean estimate than for our variance estimate. ## Special case -- IID sample We start our analysis with the benchmark case of IID data. -Consider a -sample of size $N$ generated by the following IID process, +Consider a sample of size $N$ generated by the following IID process, $$ X_i \sim \mathcal{N}(\mu, \sigma^2) @@ -1173,37 +1020,27 @@ $$ \text{MSE}(S_N, \sigma^2) = \frac{2\sigma^4}{N-1} $$ -Both estimators are unbiased and hence the MSEs reflect the -corresponding variances of the estimators. +Both estimators are unbiased and hence the MSEs reflect the corresponding variances of the estimators. -Furthermore, both MSEs are -$o(1)$ with a (multiplicative) factor of difference in their rates -of convergence: +Furthermore, both MSEs are $o(1)$ with a (multiplicative) factor of difference in their rates of convergence: $$ -\frac{\text{MSE}(S_N, \sigma^2)}{\text{MSE}(\bar X_N, \mu)} = \frac{N2\sigma^2}{N-1} \quad \underset{N \to \infty}{\to} \quad 2\sigma^2 +\frac{\text{MSE}(S_N, \sigma^2)}{\text{MSE}(\bar X_N, \mu)} = \frac{2\sigma^2 N}{N-1} \quad \underset{N \to \infty}{\to} \quad 2\sigma^2 $$ -We are interested in how this (asymptotic) relative rate of convergence -changes as increasing sampling frequency puts dependence into the data. +We are interested in how this (asymptotic) relative rate of convergence changes as increasing sampling frequency puts dependence into the data. ## Dependence and sampling frequency -To investigate how sampling frequency affects relative rates of -convergence, we assume that the data are generated by a mean-reverting -continuous time process of the form +To investigate how sampling frequency affects relative rates of convergence, we assume that the data are generated by a mean-reverting continuous time process of the form $$ dX_t = -\kappa (X_t -\mu)dt + \sigma dW_t\quad\quad $$ -where $\mu$is the unconditional mean, $\kappa > 0$ is a -persistence parameter, and $\{W_t\}$ is a standardized Brownian -motion. +where $\mu$ is the unconditional mean, $\kappa > 0$ is a persistence parameter, and $\{W_t\}$ is a standardized Brownian motion. -Observations arising from this system in particular discrete periods -$\mathcal T(h) \equiv \{nh : n \in \mathbb Z \}$with$h>0$ -can be described by the following process +Observations arising from this system in particular discrete periods $\mathcal T(h) \equiv \{nh : n \in \mathbb Z \}$ with $h>0$ can be described by the following process $$ X_{t+1} = (1 - \exp(-\kappa h))\mu + \exp(-\kappa h)X_t + \epsilon_{t, h} @@ -1215,15 +1052,13 @@ $$ \epsilon_{t, h} \sim \mathcal{N}(0, \Sigma_h) \quad \text{with}\quad \Sigma_h = \frac{\sigma^2(1-\exp(-2\kappa h))}{2\kappa} $$ -We call $h$ the *frequency* parameter, whereas $n$ -represents the number of *lags* between observations. +We call $h$ the *frequency* parameter, whereas $n$ represents the number of *lags* between observations. + +Strictly speaking, $h$ is the sampling interval, so a higher sampling frequency corresponds to a smaller $h$. -Hence, the effective distance between two observations $X_t$ and -$X_{t+n}$ in the discrete time notation is equal -to $h\cdot n$ in terms of the underlying continuous time process. +Hence, the effective distance between two observations $X_t$ and $X_{t+n}$ in the discrete time notation is equal to $h\cdot n$ in terms of the underlying continuous time process. -Straightforward calculations show that the autocorrelation function for -the stochastic process $\{X_{t}\}_{t\in \mathcal T(h)}$ is +Straightforward calculations show that the autocorrelation function for the stochastic process $\{X_{t}\}_{t\in \mathcal T(h)}$ is $$ \Gamma_h(n) \equiv \text{corr}(X_{t + h n}, X_t) = \exp(-\kappa h n) @@ -1232,19 +1067,16 @@ $$ and the auto-covariance function is $$ -\gamma_h(n) \equiv \text{cov}(X_{t + h n}, X_t) = \frac{\exp(-\kappa h n)\sigma^2}{2\kappa} . +\gamma_h(n) \equiv \text{cov}(X_{t + h n}, X_t) = \frac{\exp(-\kappa h n)\sigma^2}{2\kappa} $$ -It follows that if $n=0$, the unconditional variance is given -by $\gamma_h(0) = \frac{\sigma^2}{2\kappa}$ irrespective of the -sampling frequency. +It follows that if $n=0$, the unconditional variance is given by $\gamma_h(0) = \frac{\sigma^2}{2\kappa}$ irrespective of the sampling frequency. -The following figure illustrates how the dependence between the -observations is related to the sampling frequency +The following figure illustrates how the dependence between the observations is related to the sampling frequency -- For any given $h$, the autocorrelation converges to zero as we increase the distance -- $n$-- between the observations. This represents the "weak dependence" of the $X$ process. +- For any given $h$, the autocorrelation converges to zero as we increase the distance $n$ between the observations. This represents the "weak dependence" of the $X$ process. -- Moreover, for a fixed lag length, $n$, the dependence vanishes as the sampling frequency goes to infinity. In fact, letting $h$ go to $\infty$ gives back the case of IID data. +- Moreover, for a fixed lag length, $n$, the dependence vanishes as the sampling frequency goes to zero. In fact, letting $h$ go to $\infty$ gives back the case of IID data. ```{code-cell} ipython3 μ = .0 @@ -1272,13 +1104,11 @@ plt.show() ## Frequency and the mean estimator -Consider again the AR(1) process generated by discrete sampling with -frequency $h$. Assume that we have a sample of size $N$ and -we would like to estimate the unconditional mean -- in our case the true -mean is $\mu$. +Consider again the AR(1) process generated by discrete sampling with frequency $h$. + +Assume that we have a sample of size $N$ and we would like to estimate the unconditional mean -- in our case the true mean is $\mu$. -Again, the sample average is an unbiased estimator of the unconditional -mean +Again, the sample average is an unbiased estimator of the unconditional mean $$ \mathbb{E}[\bar X_N] = \frac{1}{N}\sum_{i = 1}^N \mathbb{E}[X_i] = \mathbb{E}[X_0] = \mu @@ -1295,55 +1125,57 @@ $$ \end{aligned} $$ -It is explicit in the above equation that time dependence in the data -inflates the variance of the mean estimator through the covariance -terms. +It is explicit in the above equation that time dependence in the data inflates the variance of the mean estimator through the covariance terms. -Moreover, as we can see, a higher sampling frequency---smaller -$h$---makes all the covariance terms larger, everything else being -fixed. +Moreover, as we can see, a higher sampling frequency---smaller $h$---makes all the covariance terms larger, everything else being fixed. -This implies a relatively slower rate of convergence of the -sample average for high-frequency data. +This implies a relatively slower rate of convergence of the sample average for high-frequency data. -Intuitively, stronger dependence across observations for high-frequency data reduces the -"information content" of each observation relative to the IID case. +Intuitively, stronger dependence across observations for high-frequency data reduces the "information content" of each observation relative to the IID case. -We can upper bound the variance term in the following way +We can upper bound the variance term in the following way, where $\rho \equiv \exp(-\kappa h)$ and $\gamma(0) = \frac{\sigma^2}{2\kappa}$ is the unconditional variance $$ \begin{aligned} -\mathbb{V}(\bar X_N) &= \frac{1}{N^2} \left( N \sigma^2 + 2 \sum_{i=1}^{N-1} i \cdot \exp(-\kappa h (N - i)) \sigma^2 \right) \\ -&\leq \frac{\sigma^2}{2\kappa N} \left(1 + 2 \sum_{i=1}^{N-1} \cdot \exp(-\kappa h (i)) \right) \\ -&= \underbrace{\frac{\sigma^2}{2\kappa N}}_{\text{IID case}} \left(1 + 2 \frac{1 - \exp(-\kappa h)^{N-1}}{1 - \exp(-\kappa h)} \right) +\mathbb{V}(\bar X_N) &= \frac{\gamma(0)}{N^2} \left( N + 2 \sum_{i=1}^{N-1} i \cdot \rho^{N - i} \right) \\ +&\leq \frac{\gamma(0)}{N} \left(1 + 2 \sum_{i=1}^{N-1} \rho^{i} \right) \\ +&= \underbrace{\frac{\sigma^2}{2\kappa N}}_{\text{IID case}} \left(1 + 2 \rho \, \frac{1 - \rho^{N-1}}{1 - \rho} \right) \end{aligned} $$ -Asymptotically, the term $\exp(-\kappa h)^{N-1}$ vanishes and the -dependence in the data inflates the benchmark IID variance by a factor -of +Asymptotically, the term $\rho^{N-1}$ vanishes and the dependence in the data inflates the benchmark IID variance by a factor of $$ -\left(1 + 2 \frac{1}{1 - \exp(-\kappa h)} \right) +\left(1 + \frac{2 \rho}{1 - \rho} \right) = \frac{1 + \exp(-\kappa h)}{1 - \exp(-\kappa h)} $$ -This long run factor is larger the higher is the frequency (the smaller -is $h$). +In fact, $N \, \mathbb{V}(\bar X_N)$ converges to $\gamma(0) \frac{1 + \rho}{1 - \rho}$, so the bound is attained in the limit. + +This long run factor is larger the higher is the frequency (the smaller is $h$). + +Therefore, we expect the asymptotic relative MSEs, $B$, to change with time-dependent data. -Therefore, we expect the asymptotic relative MSEs, $B$, to change -with time-dependent data. We just saw that the mean estimator's rate is -roughly changing by a factor of +We just saw that the mean estimator's rate is roughly changing by a factor of $$ -\left(1 + 2 \frac{1}{1 - \exp(-\kappa h)} \right) +\left(1 + \frac{2 \rho}{1 - \rho} \right) = \frac{1 + \exp(-\kappa h)}{1 - \exp(-\kappa h)} $$ -Unfortunately, the variance estimator's MSE is harder to derive. +The variance estimator's asymptotic MSE can also be computed in closed form. + +Because the process is Gaussian, Bartlett's formula implies that $N \, \mathbb{V}(S_N)$ converges to $2 \sum_{k=-\infty}^{\infty} \gamma_h(k)^2 = 2 \gamma(0)^2 \frac{1 + \rho^2}{1 - \rho^2}$, while the squared bias of $S_N$ is of order $N^{-2}$. + +Dividing by the limit of $N \, \mathbb{V}(\bar X_N)$ gives + +```{math} +:label: B-closed-form + +B(h) = 2 \gamma(0) \, \frac{1 + \rho^2}{(1 + \rho)^2}, \qquad \rho = \exp(-\kappa h) +``` + +As $h \to \infty$, $\rho \to 0$ and $B(h)$ approaches the IID benchmark $2 \gamma(0)$, while as $h \to 0$, $\rho \to 1$ and $B(h)$ falls to $\gamma(0)$. -Nonetheless, we can approximate it by using (large sample) simulations, -thus getting an idea about how the asymptotic relative MSEs changes in -the sampling frequency $h$ relative to the IID case that we -compute in closed form. +We can confirm this with (large sample) simulations, which show how the asymptotic relative MSE changes with the sampling frequency $h$ relative to the IID case that we compute in closed form. ```{code-cell} ipython3 @jit @@ -1367,7 +1199,14 @@ def sample_generator(h, N, M): ``` ```{code-cell} ipython3 +# sample_generator is compiled by Numba, whose random number generator +# must be seeded from inside a jitted function +@jit +def set_seed(seed): + np.random.seed(seed) + # Generate large sample for different frequencies +set_seed(1234) N_app, M_app = 1000, 30000 # Sample size, number of simulations h_grid = np.linspace(.1, 80, 30) @@ -1398,15 +1237,14 @@ fig, ax = plt.subplots(figsize=(8, 5)) ax.plot(h_grid, rate_h, c='darkblue', lw=2, label=r'large sample relative MSE, $B(h)$') ax.axhline(benchmark_rate, c='k', ls='--', label=r'IID benchmark') -ax.set_title('Relative MSE for large samples as a function of sampling \ - frequency \n MSE($S_N$) relative to MSE($\\bar X_N$)') +ax.set_title('Relative MSE for large samples as a function of sampling frequency\n' + 'MSE($S_N$) relative to MSE($\\bar X_N$)') ax.set_xlabel('Sampling frequency, $h$') ax.legend() plt.show() ``` -The above figure illustrates the relationship between the asymptotic -relative MSEs and the sampling frequency +The above figure illustrates the relationship between the asymptotic relative MSEs and the sampling frequency - We can see that with low-frequency data -- large values of $h$ -- the ratio of asymptotic rates approaches the IID case. @@ -1416,3 +1254,328 @@ relative MSEs and the sampling frequency dependence gets more pronounced, the rate of convergence of the mean estimator's MSE deteriorates more than that of the variance estimator. + +This experiment holds the number of observations $N$ fixed, so a higher sampling frequency also shortens the calendar span of the sample. + +If the span is held fixed instead, sampling more often does not improve the estimate of the drift at all, while it does sharpen the estimate of the variance {cite}`merton1980estimating`; see Exercise {ref}`bl_ex3`. + +## Exercises + +```{exercise} +:label: bl_ex1 + +**How much does $\tau$ matter, and what value of $\tau$ would a Bayesian choose?** + +Return to the 10-asset example at the start of the lecture (re-create its data with `np.random.seed(12)`, since later cells overwrite the variable names). + +Use the market-implied risk aversion $\delta_m$ both to construct $\mu_{BL}$ and to compute portfolios. + +1. Show analytically that $\tilde \mu = (\tau \mu_{BL} + \hat \mu)/(1+\tau)$, and hence that $\tilde w = (\tau w_m + w)/(1+\tau)$, where $w = (\delta_m \hat\Sigma)^{-1}\hat\mu$ is the mean-variance portfolio. Confirm this numerically on a grid of $\tau \in [10^{-3}, 10^3]$. +1. Plot $\|\tilde w - w_m\|$ and $\|\tilde w - w\|$ against $\tau$ on a log scale. +1. In the Bayesian interpretation of the lecture, $\hat\mu \mid \mu \sim \mathcal N(\mu, \tau\Sigma)$. If $\hat \mu$ is the sample mean of $T$ i.i.d. draws, what value of $\tau$ does that imply? Compute the implied portfolio and describe it. +1. Now put the scalar on the prior instead, as in the Black-Litterman calibration of He and Litterman: $\mu \sim \mathcal N(\mu_{BL}, \tau_p \Sigma)$ and $\hat\mu \mid \mu \sim \mathcal N(\mu, \Sigma/T)$. Compute the weight on the market portfolio and the extreme portfolio weights for $\tau_p \in \{0.01, 0.05, 0.25\}$. +``` + +```{solution-start} bl_ex1 +:class: dropdown +``` + +Because both covariance matrices are proportional to $\Sigma$, + +$$ +\tilde \mu = \left( (1 + \tau^{-1}) \Sigma^{-1} \right)^{-1} \Sigma^{-1} \left( \mu_{BL} + \tau^{-1} \hat \mu \right) += \frac{\tau \mu_{BL} + \hat \mu}{1 + \tau} +$$ + +Premultiplying by $(\delta_m \Sigma)^{-1}$ and using $(\delta_m \Sigma)^{-1} \mu_{BL} = w_m$ gives + +$$ +\tilde w = \frac{\tau w_m + w}{1 + \tau} +$$ + +So in this version of the model the Black-Litterman portfolio is simply a convex combination of the market portfolio and the mean-variance portfolio, with weight $\tau/(1+\tau)$ on the market. + +A **large** $\tau$ expresses **little** confidence in the investor's estimate $\hat \mu$. + +```{code-cell} ipython3 +# Re-create the 10-asset example of the lecture (later cells overwrite these names) +np.random.seed(12) +N, T = 10, 200 +w_m = np.random.rand(N) +w_m = w_m / w_m.sum() +μ = (np.random.randn(N) + 5) / 100 +S = np.random.randn(N, N) +V = S @ S.T +Σ = V * (w_m @ μ)**2 / (w_m @ V @ w_m) +δ = 1 / np.sqrt(w_m @ Σ @ w_m) +sample = stat.multivariate_normal(μ, Σ).rvs(T) +μ_est = sample.mean(0).reshape(N, 1) +Σ_est = np.cov(sample.T) +d_m = (w_m @ μ_est) / (w_m @ Σ_est @ w_m) +μ_m = (d_m * Σ_est @ w_m).reshape(N, 1) +w = np.linalg.solve(d_m * Σ_est, μ_est) # mean-variance weights + +τ_grid = np.logspace(-3, 3, 200) +dist_m, dist_mv, gap = [], [], [] +for τ in τ_grid: + μ_tilde = black_litterman(1, μ_m, μ_est, Σ_est, τ * Σ_est) + w_tilde = np.linalg.solve(d_m * Σ_est, μ_tilde) + w_check = (τ * w_m.reshape(N, 1) + w) / (1 + τ) + gap.append(np.max(np.abs(w_tilde - w_check))) + dist_m.append(np.linalg.norm(w_tilde.flatten() - w_m)) + dist_mv.append(np.linalg.norm(w_tilde - w)) + +print(f"max |w_tilde - (τ w_m + w)/(1+τ)| over grid: {max(gap):.2e}") +print(f"||w - w_m|| = {np.linalg.norm(w.flatten() - w_m):.3f}") + +fig, ax = plt.subplots(figsize=(8, 5)) +ax.semilogx(τ_grid, dist_m, lw=2, label=r'$\|\tilde w - w_m\|$') +ax.semilogx(τ_grid, dist_mv, lw=2, label=r'$\|\tilde w - w\|$') +ax.axvline(1 / T, c='k', ls='--', label=r'$\tau = 1/T$') +ax.set_xlabel(r'$\tau$') +ax.legend() +plt.show() + +# Bayesian calibration: sampling covariance of μ̂ is Σ/T, so τ = 1/T +τ_T = 1 / T +w_T = np.linalg.solve(d_m * Σ_est, + black_litterman(1, μ_m, μ_est, Σ_est, τ_T * Σ_est)) +print(f"τ = 1/T: weight on w_m = {τ_T / (1 + τ_T):.4f}, " + f"min weight = {w_T.min():.3f}, max weight = {w_T.max():.3f}") + +# He-Litterman placement: prior μ ~ N(μ_m, τ_p Σ), data μ̂ ~ N(μ, Σ/T) +for τ_p in [0.01, 0.05, 0.25]: + μ_HL = black_litterman(1, μ_m, μ_est, τ_p * Σ_est, Σ_est / T) + w_HL = np.linalg.solve(d_m * Σ_est, μ_HL) + s = 1 / (1 + τ_p * T) + print(f"τ_p = {τ_p:5.2f}: weight on w_m = {s:.3f}, " + f"min weight = {w_HL.min():.3f}, max weight = {w_HL.max():.3f}") +``` + +The numerical check confirms the convex-combination formula to machine precision (about $10^{-15}$). + +The two distance curves cross at $\tau = 1$, where $\tilde w$ lies halfway between $w$ and $w_m$ (which are $4.44$ apart). + +The Bayesian calibration in part 3 is sobering. + +If $\hat\mu$ is a sample mean of $T = 200$ observations, its sampling covariance is $\Sigma/T$, so $\tau = 1/T = 0.005$. + +The posterior then puts weight $0.005$ on the market view, and the recommended portfolio has weights running from about $-1.08$ to $2.97$, essentially the extreme mean-variance portfolio. + +Taken literally, the Bayesian interpretation with a prior covariance equal to the covariance of returns lets the data swamp the prior. + +That prior is extremely diffuse: it says that uncertainty about the *mean* excess return is as large as the volatility of a single period's return. + +Part 4 shows that what matters for the posterior is the ratio of prior precision to data precision, $\tau_p T$. + +The weight on the market portfolio is $1/(1+\tau_p T)$: $0.333$, $0.091$ and $0.020$ for $\tau_p = 0.01, 0.05, 0.25$. + +The most negative weight is correspondingly $-0.72$, $-0.99$ and $-1.07$. + +So even a tight prior around market-implied returns gets overwhelmed by $200$ observations *if* the investor regards the sample mean as a view with sampling covariance $\Sigma/T$. + +In practice Black-Litterman users make their views much less precise than $\Sigma/T$; that choice, and not the Bayesian arithmetic, is what keeps the recommended portfolio near the market. + +```{solution-end} +``` + +```{exercise} +:label: bl_ex2 + +**The robust portfolio, its worst-case mean, and short positions.** + +Use the 10-asset data from the start of the lecture and set $\theta = 0.05$, so that $\gamma = \theta^{-1} = 20$. + +1. For the mean-variance portfolio $w$, minimize $w'(\hat\mu + C v) + \frac{\theta}{2} v'v$ numerically over the mean distortion $v$, where $CC' = \hat\Sigma$. Check that the minimizer is $v = -\theta^{-1} C' w$ and that the minimized value equals $w'\hat\mu - \frac{1}{2\theta} w'\hat\Sigma w$. +1. Maximize ${\sf T}[w'(\vec r - r_f {\bf 1})] - \frac{\delta}{2} w'\hat\Sigma w$ numerically and confirm that the maximizer is $w_{\rm rob} = (\delta + \gamma)^{-1} \hat\Sigma^{-1} \hat\mu$. +1. Let $\mu_{wc} = \hat\mu - \gamma \hat\Sigma w_{\rm rob}$ be the worst-case mean *at the robust optimum*. Show that a non-robust investor with risk aversion $\delta$ who believes $\mu_{wc}$ chooses $w_{\rm rob}$. In what sense is $\mu_{wc}$ like the Black-Litterman $\mu_{BL}$? +1. Compare the sign patterns of the mean-variance, robust, and Black-Litterman ($\tau = 1$) portfolios. Does robustness cure extreme long-short positions? +``` + +```{solution-start} bl_ex2 +:class: dropdown +``` + +For part 3, note that + +$$ +(\delta \Sigma)^{-1} (\mu - \gamma \Sigma w_{\rm rob}) += \frac{\delta + \gamma}{\delta} w_{\rm rob} - \frac{\gamma}{\delta} w_{\rm rob} += w_{\rm rob} +$$ + +The worst-case mean is therefore an "implied" mean in the same sense as $\mu_{BL}$: it is the belief about mean excess returns that rationalizes the recommended portfolio for an ordinary mean-variance investor. + +The difference is where the belief comes from. + +Black and Litterman choose the belief so that it supports the *market* portfolio, while the robust investor's belief comes from a malevolent agent's response to the investor's own portfolio. + +```{code-cell} ipython3 +from scipy.optimize import minimize + +# Re-create the 10-asset example (same seed as the lecture) +np.random.seed(12) +N, T = 10, 200 +w_m = np.random.rand(N) +w_m = w_m / w_m.sum() +μ = (np.random.randn(N) + 5) / 100 +S = np.random.randn(N, N) +V = S @ S.T +Σ = V * (w_m @ μ)**2 / (w_m @ V @ w_m) +δ = 1 / np.sqrt(w_m @ Σ @ w_m) +sample = stat.multivariate_normal(μ, Σ).rvs(T) +μ_est = sample.mean(0) +Σ_est = np.cov(sample.T) +C = np.linalg.cholesky(Σ_est) + +θ = 0.05 +γ = 1 / θ + +# (1) Worst-case mean shift for a given portfolio +w0 = np.linalg.solve(δ * Σ_est, μ_est) +obj = lambda v: w0 @ (μ_est + C @ v) + θ * v @ v / 2 +v_star = minimize(obj, np.zeros(N), method='BFGS', options={'gtol': 1e-12}).x +print("max |v* - (-C'w/θ)|:", np.max(np.abs(v_star + C.T @ w0 / θ))) +print("T numeric:", obj(v_star), + " closed form:", w0 @ μ_est - w0 @ Σ_est @ w0 / (2 * θ)) + +# (2) Robust portfolio and the worst-case mean as an "implied" mean +w_mv = np.linalg.solve(δ * Σ_est, μ_est) +w_rob = np.linalg.solve((δ + γ) * Σ_est, μ_est) +neg = lambda w: -(w @ μ_est - (δ + γ) / 2 * w @ Σ_est @ w) +w_num = minimize(neg, np.zeros(N), method='BFGS', options={'gtol': 1e-12}).x +print("max |w_num - w_rob|:", np.max(np.abs(w_num - w_rob))) + +μ_wc = μ_est - γ * Σ_est @ w_rob +w_implied = np.linalg.solve(δ * Σ_est, μ_wc) +print("max |w(μ_wc) - w_rob|:", np.max(np.abs(w_implied - w_rob))) + +# (3) Compare with mean-variance and Black-Litterman portfolios +print("w_rob / w_mv (elementwise):", np.round(w_rob / w_mv, 4)) +print("δ/(δ+γ) =", round(δ / (δ + γ), 4)) +d_m = (w_m @ μ_est) / (w_m @ Σ_est @ w_m) +μ_m = d_m * Σ_est @ w_m +w_bl = np.linalg.solve(δ * Σ_est, + black_litterman(1, μ_m, μ_est, Σ_est, 1.0 * Σ_est)) +for name, x in [('mean-variance', w_mv), ('robust', w_rob), + ('Black-Litterman', w_bl)]: + print(f"{name:16s} short positions: {np.sum(x < 0)}, " + f"sum |w| = {np.abs(x).sum():.3f}, sum w = {x.sum():.3f}") +print("sign pattern mv :", np.sign(w_mv).astype(int)) +print("sign pattern rob:", np.sign(w_rob).astype(int)) +print("sign pattern BL :", np.sign(w_bl).astype(int)) + +# Black-Litterman short positions as confidence in μ̂ falls (τ rises) +for τ in [1, 5, 10, 50]: + w_bl_τ = np.linalg.solve(δ * Σ_est, + black_litterman(1, μ_m, μ_est, Σ_est, τ * Σ_est)) + print(f"τ = {τ:3d}: Black-Litterman short positions = {np.sum(w_bl_τ < 0)}") +``` + +The numerical minimization reproduces the worst-case distortion $v = -\theta^{-1}C'w$ (to about $4\times 10^{-8}$) and the closed form for ${\sf T}$ (to machine precision). + +The numerically maximized robust portfolio matches $(\delta+\gamma)^{-1}\hat\Sigma^{-1}\hat\mu$. + +A mean-variance investor who holds the worst-case mean $\mu_{wc}$ chooses exactly $w_{\rm rob}$ (error about $10^{-14}$). + +Every element of $w_{\rm rob}/w$ equals $0.4803 = \delta/(\delta+\gamma)$. + +Robustness of this kind is observationally equivalent to raising risk aversion from $\delta$ to $\delta + \gamma$. + +It scales down gross exposure (here $\sum_i |w_i|$ falls from $12.08$ to $5.80$) but leaves the pattern of long and short positions untouched: both portfolios short the same 5 assets. + +With $\tau = 1$ the Black-Litterman portfolio also still shorts those 5 assets in this example, but it lies on the way to $w_m$, which has no short positions. + +Raising $\tau$ removes short positions one by one (4 remain at $\tau = 5$, 2 at $\tau = 10$, none at $\tau = 50$), something no value of $\theta$ can do in the robust model. + +The Black-Litterman adjustment changes the *direction* of the portfolio, while this robust adjustment changes only its *scale*. + +```{solution-end} +``` + +```{exercise} +:label: bl_ex3 + +**Means versus variances: a closed form for $B(h)$ and a fixed-span experiment.** + +1. For the discretely sampled Ornstein-Uhlenbeck process in the appendix, let $\rho = \exp(-\kappa h)$ and $\gamma(0) = \sigma^2/(2\kappa)$. Starting from $N\,\mathbb V(\bar X_N) \to \gamma(0)\frac{1+\rho}{1-\rho}$ and Bartlett's formula $N\,\mathbb V(S_N) \to 2 \sum_{k=-\infty}^{\infty} \gamma_h(k)^2 = 2\gamma(0)^2 \frac{1+\rho^2}{1-\rho^2}$, verify the closed form {eq}`B-closed-form` for the asymptotic relative MSE $B(h)$. Plot it against the simulated `rate_h` from the lecture. Where does the simulation depart from the formula, and why? +1. The appendix holds the number of observations $N$ fixed as $h$ changes, so a higher frequency also means a shorter calendar span. Instead, hold the span fixed at $20$ years. Let log prices follow a Brownian motion with drift $m = 0.06$ and volatility $s = 0.2$, so that returns over an interval $h$ are i.i.d. $\mathcal N(m h, s^2 h)$. For annual, monthly, weekly and daily sampling, compute by simulation the relative RMSEs of the annualized drift estimator $\hat m = \sum_i r_i / 20$ and of the annualized variance estimator $\hat\sigma^2 = \sum_i (r_i - \bar r)^2/((n-1)h)$. +``` + +```{solution-start} bl_ex3 +:class: dropdown +``` + +Dividing the two asymptotic variances (the squared biases are of smaller order) gives {eq}`B-closed-form` + +$$ +B(h) = 2 \gamma(0) \, \frac{1 + \rho^2}{(1 + \rho)^2} +$$ + +As $h \to \infty$, $\rho \to 0$ and $B \to 2\gamma(0)$, the IID benchmark. + +As $h \to 0$, $\rho \to 1$ and $B \to \gamma(0)$. + +The relative MSE therefore falls by at most a factor of two as the sampling frequency rises with $N$ fixed. + +```{code-cell} ipython3 +# Part 1: closed-form asymptotic relative MSE +ρ_grid = np.exp(-κ * h_grid) +B_asym = 2 * var_uncond * (1 + ρ_grid**2) / (1 + ρ_grid)**2 + +fig, ax = plt.subplots(figsize=(8, 5)) +ax.plot(h_grid, rate_h, lw=2, label='simulated $B(h)$, $N=1000$') +ax.plot(h_grid, B_asym, 'k--', lw=2, label='asymptotic formula') +ax.axhline(var_uncond, c='r', ls=':', label=r'limit as $h \to 0$') +ax.set_xlabel('sampling interval $h$') +ax.legend() +plt.show() + +for k in [0, 1, 5, 29]: + print(f"h = {h_grid[k]:6.2f}: simulated B = {rate_h[k]:.3f}, " + f"asymptotic B = {B_asym[k]:.3f}, " + f"N·MSE(mean)/γ(0) = {N_app * mse_mean[k] / var_uncond:7.2f}") + +# Part 2: Merton's fixed-span experiment +np.random.seed(1234) +m, s = 0.06, 0.20 # annual drift and volatility of log price +span = 20 # years of data +M = 200_000 # replications +for label, h in [('annual', 1), ('monthly', 1/12), + ('weekly', 1/52), ('daily', 1/252)]: + n = int(round(span / h)) + # returns r_i ~ N(m h, s^2 h), i = 1, ..., n; use exact sampling distributions + sum_r = np.random.normal(m * h * n, s * np.sqrt(h * n), M) + ss = s**2 * h * np.random.chisquare(n - 1, M) # sum of squared deviations + m_hat = sum_r / span + s2_hat = ss / ((n - 1) * h) + rmse_m = np.sqrt(np.mean((m_hat - m)**2)) + rmse_s2 = np.sqrt(np.mean((s2_hat - s**2)**2)) + print(f"{label:8s} n = {n:5d}: RMSE(m̂)/m = {rmse_m / m:.3f}, " + f"RMSE(σ̂²)/σ² = {rmse_s2 / s**2:.4f}") +``` + +The closed form tracks the simulation closely except at the very highest frequencies. + +Take $h = 0.1$, where $1/(1-\rho) \approx 100$: a sample of $N = 1000$ spans only about ten "correlation times", so the asymptotic formula ($1.25$) overstates the finite-sample ratio (about $1.07$ in the simulation). + +The printed values of $N \cdot \text{MSE}(\bar X_N)/\gamma(0)$ show what the ratio hides: going from $h = 80$ to $h = 0.1$ inflates the MSE of the mean by a factor of roughly $180$. + +The MSE of the variance estimator also deteriorates, by a factor of roughly $80$, so their ratio falls only from about $2.5$ to about $1.1$. + +Part 2 isolates the economically relevant comparison. + +With the calendar span fixed at $20$ years, the relative RMSE of the drift estimator is $0.75$ at every sampling frequency. + +That matches $s/(m\sqrt{20}) = 0.745$, because $\hat m = (\log P_{20} - \log P_0)/20$ depends only on the endpoints. + +The relative RMSE of the variance estimator falls from $0.33$ (annual) to $0.09$ (monthly), $0.04$ (weekly) and $0.02$ (daily), in line with $\sqrt{2/(n-1)}$. + +This is the point made by {cite:t}`merton1980estimating`. + +Sampling more finely within a fixed span is useless for learning about mean returns but very informative about variances and covariances. + +That asymmetry is what licenses Black and Litterman, and robust decision makers, to distrust estimated means while trusting estimated covariances. + +```{solution-end} +```