Digital Strategy and Markets · Week 4

The force that shifts: the value of the future when rivals meet again sooner

Two agencies play the commission game of the slides year after year. Choose a rule for each, set how much they value next year's profit, and see which rule earns most against the rival's, and at which discount factor the answer changes.

Firm A: profit per year, discounted
 
Firm B: profit per year, discounted
 
Is the pair of strategies stable?
 
Firm A's best rule against firm B's actual play
 
Firm B's best rule against firm A's actual play
 
Smallest δ at which keeping High beats cutting
 
The setup. Two agencies play the commission game of the slides every year: both High earns 10 each, both Low earns 6 each, and an agency that cuts to Low while the rival keeps High earns 14 against the rival's 4. Choose a rule for each firm. Always High and always Low ignore the history. Tit-for-tat copies the rival's action of the previous year. Grim trigger plays High until anyone has played Low, then Low forever. Finite punishment plays Low for T years after any cut on the cooperative path, then returns to High. A firm values a profit next year at δ times a profit today, and its present value is the discounted sum over the years. The page reports it per year, as (1 − δ) times the present value, so the numbers stay between 4 and 14 whatever δ is. The first chart plays the two rules against each other, and a click on any of its rectangles forces that firm's action in that year, which is the forced deviation of the paper's Figure 4 applied to the classical rules; the second chart draws the value of each of the five clean rules for a firm as δ changes, against the rival as it actually behaves, with the firm's actual play as a dark line; the third and fourth rank the same five rules and the actual play of each firm at the current δ. Without forced actions the actual play sits on the chosen rule's line.
What to notice. Start with the first preset, a cut against grim trigger at δ = 0.80. Firm A earns 14 once and 6 ever after, 7.6 per year, against the 10 it would earn by keeping High. The second chart shows why: the always Low line crosses the always High line at δ = 0.5, which is the board calculation of the slides. Drag δ below 0.5 and cutting becomes A's best option. Against finite punishment with T = 3 the crossing moves to 0.54, and with T = 1 the lines never cross: a one-year punishment costs the cutter 4 next year for a gain of 4 today, and no δ below 1 makes that a loss. Then the mistakes. With two grim triggers and a 5 percent chance of an accidental cut, cooperation ends for good within a few years, and both firms average about 8 per year, where steady cooperation gives 10. Two finite punishments lose a few years after each accident and recover. Two tit-for-tats echo an accident back and forth, one firm High while the other is Low, until a second accident ends the echo. This is the reason the programs in Calvano et al. learn a punishment that ends: they keep experimenting, and a rule that never forgives would ruin both.
How this maps to the slides and the paper. The board calculation on the slide about future losses compares 10/(1 − δ) with 14 + 6δ/(1 − δ) and gives δ ≥ 1/2. The first preset is that comparison, and the crossing in the second chart is its picture. The slide on classical strategies names the three history-dependent rules used here. Calvano, Calzolari, Denicolò and Pastorello (2020) find that their programs learn none of these rules exactly: the punishment they learn lasts a few periods and prices recover gradually, which sits between tit-for-tat and finite punishment on this page, and a grim trigger never appears in their sessions. Axelrod's tournaments made tit-for-tat famous, and the echo in the last preset is its known weakness when players make mistakes.
Limits. The menu has five rules, so "best" means best among them. Against tit-for-tat, for instance, alternating Low and High beats every rule on the menu when δ is below about 0.68. Finite punishment here is a public phase: both firms count the same T years after any cut, including an accidental one of their own, which is how the rule is written in the theory. Mistakes hit both firms at the same rate. The game has two prices, where the Q-learning page has fifteen, and the rules here are given to the firms, where the programs on that page learn theirs.