@gabepereyra@harvey Does the firm do the work itself (upload data, kick off fine tuning from Tenet), or does Harvey run the engagement end to end, selecting and cleaning the firm’s data, training the model, evaluating it, and serving it?
Have you thought about building an alignment dataset designed for modern situations? reward hacking break rules because it's designed on the framework “ends justify the means”...
The goal would be to give models examples of the underlying ethical reasoning rather than just labeling things as safe or unsafe.
Curious what you think about this approach.
reminds me of ethics. knowing the difference between right and wrong is an important function of society, but i think the subject has been hijacked by people who are more interested in demonstrating that they hold the “right” moral beliefs than actually reasoning through what is right and wrong. ethics becomes more about signaling virtue than pursuing truth.