Week one: define the problem
Choose a repeated task, decision owner and small user group. Measure current time and errors, define allowed data and success criteria, and identify actions that remain manual. Set this scope before selecting tools.
Weeks two and three: limited use
Build a test set and an end-to-end prototype. Measure correction time as well as generation time. Add newly discovered edge cases and review failures regularly. Use de-identified data where possible and avoid broadening permissions as a shortcut.
Week four: decide
Compare quality, time and total cost against the baseline. Choose conditional expansion, revision or stopping. Preserve lessons even when the pilot ends: discovering poor fit early can prevent an expensive rollout.