본문 바로가기

전체 글

(383)
[ AI/Regression ] X와 Y 사이에 어떤 관계가 있을까?(Feat. 선형회귀) 회귀선이 평균보다 훨씬 잘 맞는다면, X가 Y 예측에 실제로 도움이 되는가?회귀는 X를 보고 Y를 예측하는 선을 만들고, 그 선이 실제값을 얼마나 잘 맞추는지 보는 방법 10, 20, 30이라는 숫자에서 아무 정보도 없이 모두에게 같은 값을 예측 한다면10 => 오차 큼30 => 오차큼20 => 가장 가운데 값평균은 회귀모델과 비교하기 위한 기준선 전체 변화 = 모델이 설명한 것 + 모델이 못 설명한 것SST: 실제값이 평균에서 얼마나 떨어져 있나 → 전체 변화량SSR: 평균보다 회귀선이 얼마나 설명했나 → 모델이 설명한 부분SSE: 회귀선과 실제값이 얼마나 떨어졌나 → 모델이 틀린 부분 SSE:모델이 틀린 양을 전부 모아놓은 점수SST:실제 데이터들이 평균 주변에 얼마난 넓게 흩어져 있는가SSR:..
[ LLM/Transformer Block ] Transformer Block 안을 들여다 보자 Trasnformer Block에 대해서 좀 더 자세히 알아 보자1. Prompt 입력2. Tokenization- 텍스트 -> token ID3. Embedding- ID -> 숫자 벡터4. Transformer Blocks- 문맥 계산 x 여러 층5. Final Norm- 최종 scale 정돈6. LM head- 벡터 -> Vocabulary 점수7. Logits / 확률- 후보 token의 점수8. Next Token- 하나 선택-> 다시 입력Attention 다른 Token에서 무엇을 가져올까를 계산 MLP현재 Otken의 표현을 어떻게 바꿀까를 계산 RMSNorm계산들이 아용할 숫자의 전체 scale을 정돈 Redisudal기존 정보를 유지하면서 새 계산 결과를 더함 Transformer B..
[ LLM/Ingerence ] 추론에 대해서 간단히 알아보 사용자가 질문을 주면, 모델이 그 뒤에 올 단어(Token)을 계속 생성하는 것. TrainingInference목적모델을 학습모델을 사용해서 답변을입력대량의 학습 데이터사용자의 PromptWeight 변경OX결과더 나은 모델답변 token예시GPT를 만드는 과정Chat GPT에게 질문 Training1. 수많은 문장2. 모델이 예측3. 정답과 비교4. 오차 계산5. weight 수정 Inference │1. Prefill │ 질문을 읽는다 2. Decode 답을 쓴다Prefill = "질문을 읽는 단계"Decode = "답변을 쓰는 단계" Decode는 한 번에 한 번씩 밖에 출력 안됨.1. token 하나를 만들 때 처리해야 하는 모델의 비용을 줄이자 Cheaper Model == Token ..
[ AI/Kimi ] 중국이 처음 미국을 이겼다. 항목k2/k2.6k3의미전체 파라미터약 1조2.8조저장된 지식, 기능 후보 공간 확대전문가수384개 중 8개 선택896개 중 16개 선택계산량 증가를 억제하면서 모델 욕량 확대문맥 길이128K→256K(2025.9)→256K 유지1M 토큰대형 코드 저장소, 다수 문서, 장기간 작업 처리핵심 attention주로 MLAKDA+Gated MLA장문 처리 비용과 기억 효율 개선깊이 방향 연결일반 RedisudalAttention Residual이전 층의 정보를 선택적으로 재사용독립 Intelligence Index4457종합 능력의 실질적 상승 KDA 란KDA는 긴 대화를 전부 KV cache에 쌓지 않고, 새 토큰이 들어올 때마다 작은 숫자 메모리에서 오랜된 관계를 선택적으로 지우고 새로운 관계를 써 넣는..
[ AI/LLM ] 어떻게 검증할 것인가? 모델을 답변을 이진 질문으로 검증하는 방법https://arxiv.org/abs/2606.27226 Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-ImprovementEvaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, and holistic LLM judges often produce opaque scores that are hard to debug. ..
[ AI/Medical ] 2025년 12월 2개의 Medical Models과 3개의 프론티어 모델 비교(Feat. GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6.) Do you think that specialised in a certain field model is better than a frontier model? 이 글은 아래의 논문을 글을 바탕으로 작성했습니다.https://www.nature.com/articles/s41591-026-04431-5 테스트 총 3단계로 구성Our evaluation has three stages: (1) 500 MedQA questions testing medical knowledge, (2) 500 HealthBench items measuring alignment with clinicians and (3) the real clinical queries (RCQ) benchmark, built from 100 de-..
[ AI/trasnformer ] QKV와 Attention Query와 Key가 연관도를 계산한다. 이것이 함의하고 있는 의미는내적값이 크면, 모델은 “i번째 토큰이 j번째 토큰의 정보를 많이 참고해야 한다”고 판단 예시Data visualization empowers users to create ...“create”의 Query가 “users”나 “visualization”의 Key와 높은 점수를. 그러면 create 위치의 출력은 users/value, visualization/value 등을 더 많이 섞어서 만듬. 조금 더 다른 예시로 검색을Query = 검색어Key = 문서의 검색용 색인/태그Value = 실제 문서 내용 검색어와 색인이 잘 맞으면 그 문서의 내용을 가져옴여기서 “검색어와 색인의 유사도”가 Query-Key 내적이고, “가져오는 ..
[ UIUX/buttons ] 저장, 취소 그리고 삭제의 구성 디테일의 UX 버튼에 대한 생각 Button alignment responds automatically for right-to-left languages, where the confirmation button is aligned to the left edge.- Material Degisn - 버튼 분류TypemeaningFor instanceConstructive워크플로우 진행/완료Save, Submit, Continue, Publish, SendDismissive워크플로우 취소/뒤로Cancel, Close, Back, DiscardDesctructive데이터 삭제/제거Delete, Remove, Discard, ResetNeutral부가 액션Export, Duplicate, Filter, Share..
[ UI/Card ] Card UI와 상호작용(Interaction)(Feat : hover) Card UI를 만들다 이게 버튼일까? 아닐까?라는 고민 Card UI를 만들었는 데, 과연 사용자는 이것을 클릭할까?그들이 클릭을 하게 하려면 어떤 것들의 정보가 있어야 할까?그렇다면 어떤 것들이 더 우순 순위를 둬야 할까?이러한 고민에서 시작된 것 결론:순위Effect1순위hover:shadow(optional : scale)2순위Background-colorPointer는 당연히 default Norman의 Perceived Affordance / Signifier 이론Norman referred to the affordance found in screen-based interfaces as 'perceived', on the grounds that users form and develop no..
[ AI/Harness ] 당분간은 하네스 시대가 Antrophic이 배포 실수를 했다. 그들의 하네스가 한순간에 발각 됐다. 이 글은 아래의 논문을 바탕으로 작성했습니다.https://arxiv.org/abs/2604.08224 Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness EngineeringLarge language model (LLM) agents are increasingly built less by changing model weights than by reorganizing the runtime around them. Capabilities that earlier systems expected the model to recover..
[ AI/LLMs ] Mamba 모델로 승부를 (Feat. Apple) 작은 LLM을 사용하기 위한 애플의 노력 https://arxiv.org/abs/2604.14191 Attention to Mamba: A Recipe for Cross-Architecture DistillationState Space Models (SSMs) such as Mamba have become a popular alternative to Transformer models, due to their reduced memory consumption and higher throughput at generation compared to their Attention-based counterparts. On the other hand, the community haarxiv.org Transforme..
[ Edu/AI ] AI라는 툴과 교육이라는 방향성과 목적(Feat. 샤) 우리는 AI를 교육에 어떻게 접목시켜야 하는가. 이 글은 교육을 바꾸는 사람들 사이트에서 보고 추려온 내용입니다. 교육을 다시 묻는 책들(3) - “AI 시대, 어떻게 가르칠 것인가” | 교육을바꾸는사람들Priten Shah, 『AI and the Future of Education: Teaching in the Age of Artificial Intelligence』(John Wiley & Sons, 2023) 비평 AI에 관한 책들이 쏟아지고 있다. 많은 것들이 기술의 가능성을 과장되게 나열하거나, 반대로21erick.org 1956년 벤저민 블룸이 제안한 분류학의 관점으로보면, 교육의 목표는 총 6 단계로 분류기억 -> 이해 -> 적용 -> 분석 -> 평가 -> 창조아래 단계에서 위 단계로 올라갈..