전체 글

지니의 공부 아카이브
- resaerch -

(Xiaomi MiMo-VL 리뷰) MiMo-VL Technical Report

- 작성중 (초안(LLM) + a(human)https://github.com/XiaomiMiMo/MiMo-VL1. 개요냅다 성능, MMMU-Pro에 대해서는 높은척 낮음 (높긴함)MiMo-VL-7B는 Xiaomi LLM-Core 팀이 개발한 7B 파라미터 규모의 비전-언어 모델일반 시각 이해와 다중모달 추론에서 최첨단 성능을 보이며, 오픈소스로 공개 2. 모델 구성MiMo-VL-7B는 세 가지 주요 모듈(흔한 구조) - Qwen2.5-ViT + MiMo-7B-BaseViT 인코더: 네이티브 해상도를 유지해 미세한 시각 정보 보존MLP 프로젝터: 시각·언어 표현 효율적 정렬MiMo-7B 언어 모델: 복잡한 추론에 최적화된 디코더 3. 사전 학습 단계 - 4단계 학습역시나 뒤로 갈수록 고품질의 어려운 데..

- resaerch -

(Xiaomi MiMo LLM 리뷰) MiMo: Unlocking the Reasoning Potential of Language ModelFrom Pretraining to Posttraining

- 작성중 (초안(LLM) + a(human)Xiaomi MiMo LLM: Technical Analysis and Performance Reviewhttps://github.com/XiaomiMiMo/MiMo GitHub - XiaomiMiMo/MiMo: MiMo: Unlocking the Reasoning Potential of Language Model – From Pretraining to PosttrainingMiMo: Unlocking the Reasoning Potential of Language Model – From Pretraining to Posttraining - XiaomiMiMo/MiMogithub.com 1. MiMo LLM 개요샤오미는 2025년 4월 자사 최초의 초거대 ..

til/dl

LLM training 정리 - 23.08

adapter - LoRA - IA3 정리 https://devocean.sk.com/blog/techBoardDetail.do?ID=164779&boardType=techBlog LoRA 개념 https://da2so.tistory.com/79 허깅페이스 구현 LoRA https://huggingface.co/docs/diffusers/main/en/training/lora LLM training How To Train Large Language Models 대략적인 튜토리얼 https://www.datacamp.com/tutorial/how-to-train-a-llm-with-pytorch LangChain official LangChain LangChain - use case

trial and error log

[Excel to Google Form] 엑셀을 불러와서 구글 폼에서 사용하기

1. 아마존(AMT)을 사용하지 않고 주변인들에게 정성평가 할일이 생긴다면 아주 유용하게 쓰일 것 2. 구글폼을 이렇게 바꿔서 쓰면되는데 좀 비효율적으로 진행한 정성평가가 많다... 확장 프로그램은 Form builder for sheets https://www.youtube.com/watch?v=Ku1jZkU7kYI

daily

2022년 회고록

2022년 회고록이 누울자리 1학기 1월~3월 4월 4월~5월 6월 7~8월 2학기 9월~10월 11~12월 4월 5월

trial and error log

언어모델 (BERT, RoBERTa 등) 모델 저장 크기

파이토치로 모델 저장시에 model 통째로가 아니라 가중치만 저장하는데 사이즈가 너무 커져서 찾아봄 나는 주로 RoBERTa - base를 쓰고 base이기 때문에 저장시에 용량이 1GB 는 넘지 않아야 하는데 1.3GB 내외로 저장되어서 서버 용량 관리에 애먹었다. Model become 3 times larger after finetune? #63 https://github.com/google-research/bert/issues/63 Model become 3 times larger after finetune? · Issue #63 · google-research/bert A pretrained bert large model's ckpt file is about 1.3GB, after finetun..

지니티토리
지니티토리