Digital Strategy and Markets · Week 4

The force that shifts: the cost of watching rivals and answering them

Two pricing programs set prices every period, see each other's last prices, and learn from their own profits alone. This page runs the experiment of Calvano, Calzolari, Denicolò and Pastorello (2020) in the browser. Press run and watch what the programs learn, then what they do when one of them is forced to cut its price.

Bertrand-Nash price of the one-shot game
 
Monopoly price
 
Converged price (mean of the two firms)
 
Profit gain Δ: share of the gap between Nash and monopoly profit captured
 
Periods to convergence
 
Deviator's discounted profit after the forced cut, against no cut
 
The setup. Two firms sell differentiated products at the same cost and meet every period. Each firm's price is set by a program that sees the state, the two prices charged last period, and holds a table with one row per state and one column per price on a fifteen-point grid. In each period the program reads the row for the current state and charges the price with the highest value, or a random price with probability ε, which fades as the run goes on. It earns its profit and moves the value of the cell it used toward that profit plus δ times the best value in the new state. That is all it is told: no demand curve, no rival's profit, no instruction to punish. The run stops when neither program has changed its preferred price in any state for 100,000 periods, which is the paper's rule.
What to notice. At the baseline the average price starts in the middle of the grid, falls toward the Bertrand-Nash level while exploration is heavy, and climbs once exploration fades, settling between the Nash and monopoly prices. The second chart repeats the experiment of Figure 4: firm 1 is forced to its short-run best response for one period, the rival answers with a cut of its own, and both prices return to the old level within about ten periods. The cut lowers the deviator's discounted profit, so it does not pay, and the punishment is temporary. A grim trigger would never survive learning, because the programs keep exploring and an accidental cut would end cooperation for good. The third chart is the learned rule itself: one row per price the rival charged last period, and in each cell the price firm 2 charges next. Press impatient firms: with δ = 0.35 the future is worth little, the profit gain falls, and the answer to the cut flattens, since there is little to punish with. Short exploration converges in a third of the periods with a smaller gain.
How this maps to the paper. Calvano, Calzolari, Denicolò and Pastorello (2020, American Economic Review) run this game with δ = 0.95, α = 0.15, β = 4 × 10⁻⁶, a fifteen-price grid from 10 percent below the Bertrand-Nash price to 10 percent above the monopoly price, and one period of memory, over 1,000 sessions. Convergence takes from about 400,000 periods to several million, depending on how long exploration lasts. Their average profit gain Δ is 0.85. The forced cut of Figure 4 lowers the deviator's discounted profit by 3 to 4 percent (Table 3), and prices return to their old level within about ten periods. Figure 3 reports the gain falling as δ falls, to 16 percent at δ = 0.35, and footnote 26 puts the grim-trigger threshold for this demand at about 0.40, which the page computes as from the continuous game. This page runs one to ten sessions of the same algorithm under the same rules, so its numbers sit near the paper's averages without matching them.
Limits. One session is one draw of the random exploration, and sessions differ: most end at a fixed pair of prices, some at a cycle of two or more, and the response to the cut is averaged over the points of the cycle, as in the paper. The convergence rule stops the run once exploration has faded, so the cells of states that are rarely visited late in the run hold what the programs learned while they were still exploring, which is why the third chart is smooth on the converged path and rough away from it. The page caps a session at 20 million periods. The million periods that learning takes limit the transfer to real markets, and the assumption that each program observes the rival's price is the monitoring cost of the session's question.