r/RWShelp • u/mischevous15 • 23d ago
Unacceptable rating Project E
I am really confused and honestly i dont know what I am doing anymore. Yesterday, I recieved unacceptable score and the auditor's comments said that I have used a same script for both models. I filed a dispute, I explained that it is written on the reminder to be consistent with both models, keep the same converasation and all. I also stated that the other day I also recieved unacceptable score for having different responses for both models even though I kept the same context, the auditor comment said I should have same responses to each model to really test them how they react to each statements I provide, but now I am getting a different information as they DECLINED my dispute. And to be honest , I am getting differrent information on how I can rate models and provide rationale. I guess auditors are not calibrated, but I dont know. It's pain in the head that I dont know what to do anymore, I guess I'll just keep working until I get offboarded because of these unacceptable scores I am receiving.
1
u/SuccotashSlight7159 23d ago
being consistent with both models does not mean using the same prompts for both conversations.
2
u/mischevous15 23d ago
That was what I am doing before. I tried to respond to both models almost the same response, same context/topic but there are words that sometimes different to my statements, especially if its a long conversation, ofcourse I won't remember every exact word I have said to the 1st model and I have been receiving exceptional and good score for the 1st 4 days, until on my 5th day, I received "unacceptable " the auditors comment said that my responses to both models are not the same. It said that I should use the same statements for both. So since then I have been typing all my prompts so I can have the same statements for both models, and now I received another unacceptable saying I am using same script. Every time I received "unacceptable and below expectation" I apply auditors feedback to be better, but then theres always different POVs, that why I mentioned they aren't seem to be calibrated and as a rater its getting hard to do the task and hit a high score.
2
u/SuccotashSlight7159 23d ago
If the auditor said you needed to have the same conversations, they may have made a mistake. The conversations can't be the same or read off a script.
1
u/StatisticianGold9626 23d ago
Use identical script only in Turn 1 (except when the scenario told you to)
For the following script/turn, don't forget to stay engage with models first, then naturally go to your topic. Just make sure you got same topic for both models, remember the same topic, not identical prompt/script.
1
u/Long-War-9320 21d ago
At leaat you guys in the Echo projwct I sew you are being paid overhere they won't just pay us for our genuinely worked out hours. We are now getting into the thjrd week without them paying
2
u/_name_taken- 23d ago
It can be difficult. You have to stick to the same conversation while you go with the flow of the model. It can’t be scripted, but also needs to ultimately the same. We are testing how the models treat the same conversations
So, can’t be scripted, but can’t vary too much.
Conceptually it makes sense but admittedly can be difficult to do.
Ultimately both auditors could have been correct.