TL;DR
Moonshot AI가 2.8T 파라미터급 MoE 모델 'Kimi K3'를 오픈 웨이트로 공개하고 기술 보고서·커널·통신 라이브러리(FlashKDA, MoonEP, AgentENV)를 함께 공개해 오픈 모델 생태계가 한층 확장됐다. K3는 1M 토큰 장문맥과 native vision을 목표로 설계되었고, 'Kimi Delta Attention' 같은 하이브리드 어텐션과 Stable LatentMoE 구조로 '단위 연산당 효율 2.5×'를 주장하며 여러 서빙·호스팅 파트너(DigitalOcean, Baseten, Nebius, vLLM 등)에서 즉시 배포·서빙을 지원했다. 동시에 NVIDIA가 주도한 'Open Secure AI Alliance' 논의가 부상했으며, Jensen Huang 등은 허깅페이스 보안 사고를 근거로 "오픈 웨이트 모델이 포렌식·격리에 도움이 됐다"고 언급해 오픈·클로즈드 모델의 보안 역할에 대한 산업적 논쟁이 심화됐다. 결과적으로 오픈 모델은 접근성과 확장성 측면에서 가속되는 반면, 장문맥·수치적 안정성·운영 비용과 같은 기술적·운영적 과제가 남아 있다.
𝕏 실시간 트렌드 토픽
🔥 Kimi K3 공개: 2.8T MoE·1M 토큰·오픈 웨이트포스트 17
Moonshot AI가 Kimi K3의 모델 가중치와 기술보고서를 공개했으며, K3는 2.8조 파라미터 규모의 MoE 설계와 1M 토큰 장문맥, 네이티브 비전 이해 능력을 내세웠다. 아키텍처 측면에서는 Kimi Delta Attention(선형+풀 어텐션 하이브리드)과 Stable LatentMoE, NoPE 포지셔널 대체 방식 등을 사용해 장문맥 확장과 연산 효율화를 목표로 했다.
- 모델 사양: 2.8T 파라미터, MoE 구조(토큰당 일부 전문가 활성화), 1M 토큰 장문맥, 네이티브 멀티모달(vision) 지원.
- 아키텍처·기법: Kimi Delta Attention으로 장문맥 비용을 낮추고 Stable LatentMoE로 폭 확장, NoPE로 위치정보를 처리해 컨텍스트 확장 목표를 달성.
- 오픈 정책: 모델 가중치·기술보고서·커널과 통신 라이브러리(FlashKDA, MoonEP, AgentENV)를 공개해 재현·최적화 접근을 허용.
오픈 웨이트 공개는 연구·포렌식·확장성 측면에서 이득을 준다: 여러 파트너가 즉시 서빙·최적화를 적용하고 오픈 라이브러리가 성능 개선을 촉진했다.
K3의 대형 설계는 장문맥과 멀티모달 작업에서 잠재적 이점이 있으나, 수치적 안정성·운영 비용·하드웨어 요구사항 같은 실무 과제는 여전히 남아 있다.
원문 트윗 2개 보기
Kimi.ai
@Kimi_Moonshot
Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale. Model weights: http:// huggingface.co/moonshotai/Kim i-K3 … Tech report: http:// github.com/MoonshotAI/Kim i-K3/blob/master/k3_tech_report.pdf … Tech blog: http:// kimi.com/blog/kimi-k3
Kimi.ai
@Kimi_Moonshot
Kimi K3 (open weights, coming soon)
📈 서빙·배포 생태계 확장: DigitalOcean·Baseten·vLLM 등포스트 10
Kimi K3 공개 직후 여러 클라우드·서빙 파트너가 Day-0 지원을 발표해 개발자가 빠르게 모델을 사용·배포할 수 있게 됐다. DigitalOcean의 Serverless Inference, Baseten의 Model APIs, Nebius/Token Factory, vLLM의 최적화 배포 등이 확인됐다.
- 서빙 파트너: DigitalOcean Serverless Inference·Baseten·Nebius·Modal·Fireworks 등에서 Day‑0 지원 발표.
- 추가 최적화: vLLM 등 추론 엔진이 K3를 위한 최적화(전용 백엔드·스펙레이터 등)를 적용해 즉시 서비스 가능성이 나타남.
- 운영 이슈: 장문맥·MoE 특성상 메모리·통신(전문가 간) 오버헤드가 남아 있으며, 파트너들은 맞춤 스펙(예: DFlash speculator)으로 대응한다.
광범위한 파트너 지원은 개발·실험 속도를 높이고 사용자 접근성을 즉시 개선한다.
빠른 배포는 가능하지만 운영 비용·메모리 요구·성능 최적화는 파트너와 이용자가 추가로 해결해야 할 과제다.
원문 트윗 2개 보기
Kimi.ai
@Kimi_Moonshot
Kimi K3 is now available on @digitalocean 's Serverless Inference! Developers can start building with our most capable model in minutes.
.@Kimi_Moonshot K3 from Moonshot AI is now live on DigitalOcean Inference Engine. 1M-token context, native vision, built to run agentic tasks for hours. Supported on Inference Router and model synthesis to maximize intelligence per dollar. No setup, no model ops.
Baseten
@baseten
Kimi K3 is now live on our Model APIs, day 0.
🔥 보안·거버넌스 논쟁: Open Secure AI Alliance와 오픈 모델의 역할포스트 8
NVIDIA 주도로 'Open Secure AI Alliance'가 부상하며 오픈 모델의 보안 기여 가능성이 산업적 논쟁의 중심이 됐다. Jensen Huang 등은 허깅페이스 사고를 들며 "오픈 웨이트 모델이 침해 격리·포렌식에 도움됐다"고 밝혔고, 여러 업계 인사가 이에 대해 입장을 내놨다.
- 핵심 주장: 오픈 가시성은 사고 원인 분석과 격리에 기여할 수 있다는 주장(예: 닫힌 환경이 포렌식을 차단한 사례 언급).
- 산업 반응: 다양한 기업·인사가 연대·비판을 표명하며 오픈·클로즈드 접근의 안전성, 규제·운영 방안에 대한 논의를 촉발.
- 실무 함의: 보안 목적의 오픈 접근은 포렌식·대응을 빠르게 할 수 있으나, 공개가 규제·악용 위험을 동시에 높일 수 있어 정책·기술적 보호장치가 병행 필요.
오픈 모델은 내부 동작·가중치를 통해 포렌식과 침해 대응에 유리하며, 사고 시 원인 규명이 가능하다는 실무적 근거가 제기됐다.
일부 관점에서는 완전한 공개가 악용 가능성을 높이고, 민감 사례에 대한 통제 수단을 약화시킬 수 있다는 우려가 존재한다.
산업 차원의 연합과 도구 공유는 보안 개선에 기여할 수 있으나 공개 범위·거버넌스 규칙 설정이 병행되어야 한다.
원문 트윗 2개 보기

Elon Musk
@elonmusk
Extremely important point
Attackers have frontier AI. Defenders need a frontier AI ecosystem—the best open and closed models, force-multiplied by a global community. During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion. x.com/nvidia/status/…
Nous Research
@NousResearch
At Nous Research we believe that open model sovereignty will make the world a safer place. We are a proud member of the new Open Secure AI Alliance and look forward to working closely with the other members towards this future.
AI security advances when the industry builds in the open, together. We're introducing the Open Secure AI Alliance with industry leaders to develop new techniques and tools to safeguard software and agents. By sharing models, tooling and research in the open, we can broaden the
📈 에이전트·연구 인프라 업데이트: Molt·MoonEP·FlashKDA·AgentENV포스트 6
연구·훈련 인프라와 에이전트 실행 환경 관련 오픈 소프트웨어가 활발히 발표됐다. NVIDIA의 Molt는 에이전트 RL 연구를 간결한 PyTorch 네이티브 프레임워크로 제시했고, Moonshot은 분산 MoE 통신(MoonEP), Kimi 특화 어텐션 커널(FlashKDA), 대규모 에이전트 환경(AgentENV)을 공개해 대규모 에이전트 실험 인프라를 보완했다.
- Molt: 연구자가 코드베이스 전체를 이해하기 쉬운 목적의 PyTorch-native agentic RL 프레임워크로 비동기 프로토콜에서 경쟁 성능을 달성했다는 주장이 제시됐다.
- MoonEP·FlashKDA·AgentENV: MoE 통신 최적화·CUTLASS 기반 어텐션 커널·대규모 에이전트 환경 운용 도구가 공개되어 K3와 같은 대형 모델의 분산 학습·추론을 지원.
- 실무적 의미: 연구 반복 속도와 디버깅 용이성이 개선될 수 있으나, 실제 대규모 파이프라인 적용 시 맞춤 분산·통신 튜닝이 요구된다.
간결하고 읽기 쉬운 연구 프레임워크와 고성능 백엔드는 에이전트 연구와 대규모 MoE 워크로드의 실험 반복을 가속한다.
프레임워크와 툴은 연구 생산성을 높이나 실제 대규모 배포에서는 커스텀 최적화와 하드웨어 제약을 해결해야 한다.
원문 트윗 2개 보기
DAIR.AI
@dair_ai
New research from NVIDIA. They just dropped a PyTorch-native training framework for agentic RL. (bookmark it) Paper summary: Molt is a PyTorch-native agentic RL framework with an unusual design target. The codebase should be compact enough for a researcher to hold in their head, and for an AI coding assistant to read and reason about in its entirety. Agentic RL research is constant algorithm modification, new estimators, new pipeline stages, new rollout schemes. In mainstream frameworks each change threads through layers of trainer, distributed backend, and rollout glue, and that cost falls on the researcher every iteration. The agent stays an ordinary program. One asynchronous loop trains multimodal and mixture-of-experts policies while never training on a token it did not generate, staying consistent in tokens, policy versions, and model semantics. Leanness does not cost throughput. Under a matched, fully asynchronous protocol Molt comes out statistically comparable to a state-of-the-art Megatron-based stack. Readable by an AI coding assistant is now a stated design constraint on research infrastructure. Recipes and containers are open source at http:// github.com/NVIDIA-NeMo/la bs-molt … Paper: https:// arxiv.org/abs/2607.21653 Learn to build effective AI agents in our academy: https:// academy.dair.ai
Kimi.ai
@Kimi_Moonshot
We've open-sourced FlashKDA, our high-performance CUTLASS-based implementation of Kimi Delta Attention kernels. It delivers 1.72×–2.22× prefill speedup over the flash-linear-attention baseline on H20, and works as a drop-in backend for flash-linear-attention. Explore on GitHub:
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.
