@jphme Kids will let AI do their homework, just to find out that it will also replace their future jobs, because they don't have any own competencies any more. Clever strategy!
@jphme The image illustrates the situation quite well and I share your concerns about speed vs. safety a 100%. The reason for this is, that most companies think that being the first to achieve AGI or something very close, will basically control the whole worlds economy.
@theaievangelist This is a good idea in general. However, measuring the quality of TTS models is notoriously hard. MOS (mean opinion score) is the most common means of comparison but requires large scale user studies and is therefore expensive and still somehow subjective.
@nickfloats Today, I gave Dall-E 3 a try using Microsoft Bing Image Creator. It is only mediocre for photorealistic images. However, for cartoons it is really great.
@nickfloats I tried the prompt in SDXL turbo, or to be more precise in DreamShaperXL_Turbo. I also couldn't produce a parrot. Besides the few artefacts, I still like the results.
I just tested dolphin-2_2-yi-34b and Yi-34B-Chat, both in their GPTQ version from @TheBloke. File size on disk is 18.6GB for both, however GPU memory is 23.6GB for Yi Chat and 19.5GB for DolphinYi. The runtime for 110 questions was 11:14 vs. 5:04 min. Any ideas why @erhartford?
@theaievangelist@togethercompute Why did you ask for compute and vectorDBs, which are obviously not generative AI, but not about text to image like Midjourney or Stable Diffusion, which obviously is generative AI?
@erhartford The argument is valid from a certain point of view, but still it makes lives of people that invent potential harmful things too easy. Look, I've invented these cluster bombs here. Very deadly. Still, I'm not a bad guy. Somebody else needs to deal with the bad guys who use them.
@ildefons@AlphaSignalAI If it's in the context, than ChatGPT can do the reversal, but if you open a new chat and ask directly for the son or children of Mary Lee Pfeiffer, it indeed hallucinates and invents interesting things, e.g., I got "Michelle Pfeiffer" as an answer. I did not believe it at first.
@coqui_ai@huggingface Speech quality is good, but I uploaded a 16s sample of my own voice in German and tried to synthesize German text based on that: sounds not like me at all! I wouldn't call that voice cloning.
@jon_durbin I tested the version 2.1 as this is already available as GPTQ from @TheBloke. However, the performance was rather disappointing given the overwhelming good performance of airoboros 33B based on Llama v1. It was even worse than the old 13B model.
@schnellenbachj Vielleicht findet "ihr Ökonomen" ja mal eine Lösung, in der die Luft auch Privateigentum ist und jeder den Teil atmen darf, den er selber verpestet hat, statt seine Abgase in die Allmende der Allgemeinheit zu pusten. Leider würden die Superreichen dann alles aufkaufen.
@erhartford I noticed something similar for the airoboros 70B model of @jon_durbin . He seems to have tested a series of different datasets and the current best (2.1) is on par with other 70B LlaMA 2 based models, but is way better in TruthfulQA. Has anybody looked deeper into the reasons?
@jon_durbin Thanks a lot. I started right away evaluating it and it performed very well (59.1%). Took the lead in my evaluation for 13B param models because it significantly improved over the old 13B model, whereas the previous best model from Nous Hermes did not. https://t.co/gzpYG05INZ
@erhartford In my own evaluation, dolphin 13B was on par with WizardLM (53.6% correct answers on my own dataset) and therefore only slightly worse than the best 13B models (Nous Hermes and Minotaur with 56.4% each). It falls especially short in the trick questions (1/14) and math (0/7).
@csahil28 It would be really nice, if models on huggingface would have a model card to introduce the model and give a code example. Otherwise it is so much wasted time for users to figure out how excactly the model should be prompted and used, despite AutoModel doing a good job.
@vipulved@ClementDelangue In my own tests on my self-constructed dataset, RedPajama 7B instruct doesn't perform exceptionally well compared to other instruction-tuned 7B models with 35.5% correct answers, BloomZ 7B: 35.5%, Falcon 7B instruct: 40%, MPT 7B instruct: 40.9%, WizardLM 7B: 47.3%, ChatGPT: 60.9%
@elonmusk@karpathy The problem is, that there is no specification that tells you that command a leads to result b. It is like getting an interpreter that speaks a new programming language but you have to use trial & error to figure out the syntax. Hopefully, better AI will make that superfluous.