I understand that part, but the other part:
“… the camera suddenly stops and pulls out very fast to see the objects come together and form the word.”
That implies that there is some position in 3D space where you could put the camera and all the z-separated layers would appear to line up and form the word. If that’s the case, it would seem like the ones farther away from that spot would have to be larger.
In any case, I would think the easiest thing to do would be to put the camera in its final position and arrange the layers in the positions that make the final shot work. Then create your camera fly-through and have it end up at the predefined spot. Or am I missing something?
Dan