Week one: define the problem

Choose a repeated task, decision owner and small user group. Measure current time and errors, define allowed data and success criteria, and identify actions that remain manual. Set this scope before selecting tools.

Weeks two and three: limited use

Build a test set and an end-to-end prototype. Measure correction time as well as generation time. Add newly discovered edge cases and review failures regularly. Use de-identified data where possible and avoid broadening permissions as a shortcut.

Week four: decide

Compare quality, time and total cost against the baseline. Choose conditional expansion, revision or stopping. Preserve lessons even when the pilot ends: discovering poor fit early can prevent an expensive rollout.