A convincing replacement is mostly about everything around the object: the light has to fall the same way, the shadows have to land where the old ones did, and the perspective has to match the camera that took the photo. The model reads those from your image and renders the replacement inside them, which is why the results hold up at full size instead of looking pasted in.
The practical limit is description quality. "Replace the roof" gives the model room to improvise; "replace the roof with charcoal standing-seam metal" gets you the roof you actually meant. Specific beats vague every time.