The most common limit in terms of an automatic process for this technology at large (not just us) comes from having part of a shot that goes one way and another region that goes the other way and under the other, which is not uncommon when you have camera motion combined with some action. It is a problem because at some point from a frame to another you end up with large occlusions — areas of the image that are visible in a frame but not in the next one…
If you need a rule of thumb which is 90% of the time right, that would be if a pixel travels more then 5% of an image size on a frame then you are in potential trouble zone and that’s when typically the “guidance” methods we offer such as: input of mattes separating “layers” of motion basically, and the input of point tracking data (for example will help large pans) and rotosplines (for example to disambiguate for Twixtor edges of a rotating object) become your friend if you are willing to spend the time doing it.
Pierre