問題文
A classification prompt scores 84 percent with 8 examples in the request and 85 percent with 32 examples. Each request now carries 4,000 extra tokens. What is the appropriate reading of this result?
選択肢
- The returns from adding examples have flattened, so the extra cost per request is no longer buying meaningful accuracy.
- Accuracy will keep rising roughly linearly, so the team should try 128 examples next.
- The result shows that the model has memorized the 8 examples, so the team should shuffle them on every request.
- The 1-point gain proves the prompt is now optimal and needs no further work.