Let me start with a confession: I don’t live next door to a data centre. I do, however, spend a significant amount of my working life around the technology that goes inside them, designing and thinking about AI infrastructure, power, cooling, GPUs and the increasingly difficult problem of operating extremely dense compute environments efficiently. If […]

Read More

When discussing AI inference infrastructure, the word “scaling” can mean several very different things. We might need more GPUs because a model is too large for a single GPU, more copies of a model because concurrency is increasing, or better routing between those copies because repeated context is consuming unnecessary compute. Eventually, we may also […]

Read More