For any model that reaches production, the energy spent answering questions overtakes the energy spent learning to answer them within weeks. This chapter treats inference as the primary sustainability problem in applied machine learning. It sets out one accounting identity for the carbon cost of a query, shows which engineering levers move which term of that identity, and argues that the order in which you pull those levers matters more than any single technique. The chapter closes with the multi-objective view, in which quality, latency and watts define a frontier to choose from, and with the measurement discipline needed to make any of these claims credible.