https://t.co/t0ev9iMYQD
주박이 죽었길래 그냥 하나 만들었습니다.
세션 하면서 함께 들을 브금 웹사이트입니다.
자신의 플리를 공유할 수도 있습니다.
자세한 설명, 기능, 한계점 등등은 아래의 포스트를 참고해 주세요: https://t.co/hz1Sj6EakM
행복하세요~
Introducing terminal-code: VS Code inside the terminal
- VS Code compatible CLI
- works over ssh
- syncs with your terminal theme
https://t.co/j3HThPdVPw
RAG vs. CAG, clearly explained!
RAG is great, but it has a major problem:
every query hits the vector DB. even for static information that hasn't changed in months.
this is expensive, slow, and unnecessary.
Cache-Augmented Generation (CAG) fixes this by letting the model keep static information in its key-value (KV) memory, which is what the model builds internally for every token it reads.
in fact, you can combine RAG and CAG for the best of both worlds.
here's how it works:
RAG + CAG splits your knowledge into two layers.
↳ static data (policies, documentation) gets cached once in the model's KV memory
↳ dynamic data (recent updates, live documents) gets fetched via retrieval
you get faster inference, lower costs, and less repeated work.
the trick is being selective about what you cache.
only cache static, high-value knowledge that rarely changes. cache everything and you'll hit context limits. separating "cold" (cacheable) and "hot" (retrievable) data keeps this system reliable.
you can start today. OpenAI and Anthropic already support prompt caching in their APIs.
one thing to know before you scale it.
prompt caching matches on an exact prefix, byte for byte. your cached layer only gets reused when it sits at the very front of the context in the same order every time.
↳ reorder two cached policy documents and both turn into a miss
↳ cache document A alone and document B alone, then query both, and the second one misses because the model computed its cached state without ever seeing the first
in production this looks like a small fraction of your cached blocks serving almost all the hits. the rest just sits there.
the way out comes from how attention behaves. tokens attend mostly to their own local neighborhood, and only a few reach across document boundaries. CacheBlend recomputes those few and reuses everything else from the separately cached documents.
multi-document queries run two to four times faster, quality holds, and order stops mattering.
it ships in LMCache, which is fully open source.
repo: https://t.co/TXlaLLu04a
(don't forget to star 🌟)
below, i have quoted my article on KV cache management. it covers where prefix caching stops working and how a proper caching layer fixes it.
give it a read.
cheers! :)
@jakaeui 확인했습니다. 판매글 내용에서는 구매 전 제가 프로그램의 작동 원리를 도무지 파악할 수 없기도 했고 저 또한 백업에 있어 이미지 처리에 난항을 겪고 있던 터라 문의 드린 것이니 너무 민감하게 받아들이지는 않으셨으면 합니다. 감정이 상하셨다면 그 점은 유감이며 부디 별 문제 없길 바랍니다.
@jakaeui@hello_mountain3 안녕하세요, 정말 좋은 취지로 개발하신 것은 이해합니다만 룸ID를 사용해 룸 데이터(스탠딩 이미지 파일, 채팅 등등)를 추출하는 것 자체를 코코포에서는 금지하고 있습니다. 이는 서버나 네트워크에 예상치 못한 부하를 줄 수 있으므로, 삼가시는 게 좋을 것 같습니다.
문서&코드를 일부 업데이트 했습니다.
1. 코드 관련
Bash Script를 추가하여 커맨드 입력의 수를 줄였습니다. 추후 아예 스크립트 하나로 안정적으로 모든 걸 끝낼 수 있도록 하고자 합니다.
2. 문서 관련
최근에 받았던 질문들을 기반으로 문서를 살짝 수정했습니다.
기타 문의는 이메일 바랍니다.
구글 스프레드시트를 사용해서 바로 구글 문서의 티알 내용을 정돈해주는 기능을 만들었습니다.
사용 방법: 아래의 구글 스프레드시트를 사본만들기 해서 동영상처럼 사용하시면 됩니다.
조사도 바꿔줍니다.
줄바꿈도 해줍니다.
수정 배포는 시트 참고하세요.
https://t.co/s3SZf0vatX