One machine eventually runs out of disk, or CPU, or patience. Partitioning splits the dataset so each machine owns a slice. Every hard problem here comes from the same two facts: the slices are never quite even, and the boundaries never quite stay put.