TL;DR
이번 기간에는 성능·효율성 개선을 전면에 둔 발표들이 집중됐다. Google DeepMind는 Gemini 3.6 Flash와 3.5 Flash‑Lite/Cyber를 내놓아 토큰 사용량 감소와 3.5 Flash‑Lite의 약 350 output tokens/sec급 속도를 강조했고, NVIDIA는 CoreWeave·DeepInfra 측 측정 결과를 인용해 Vera Rubin NVL72가 메가와트당 토큰 처리량에서 큰 폭의 개선을 보였다고 보고했다. Poolside는 118B 총파라미터에 토큰당 8B만 활성화되는 MoE 구조의 Laguna S 2.1을 공개하며 최대 1M 토큰 컨텍스트와 공개 가중치를 제시했다. 이러한 흐름은 전력·토큰 비용을 낮추면서도 장기 컨텍스트와 에이전트형 워크플로 지원을 확대하는 방향으로 업계가 재편되는 신호로 해석되나, 일부 모델·지표는 제한적 배포 또는 외부 검증 필요성이 병존한다.
𝕏 실시간 트렌드 토픽
🔥 Gemini 3.6 Flash와 3.5 Flash‑Lite·Cyber의 제품군 확장포스트 12
Google DeepMind는 Gemini 3.6 Flash를 통해 이전 Flash 모델 대비 토큰 사용량을 줄이는 방향으로 개선을 가했고, 3.5 Flash‑Lite는 약 350 output tokens/sec급 속도와 저비용 옵션을 목표로 설계됐다. 3.5 Flash Cyber는 코드 보안 취약점 탐지에 특화되어 제한적 파트너 배포 방침이 적용됐다.
- 3.6 Flash는 3.5 Flash 대비 토큰 사용량을 줄이는 데 중점을 둔 업그레이드임이 강조됐다.
- 3.5 Flash‑Lite는 빠른 응답이 필요한 에이전트형 워크플로를 겨냥해 약 350 output tokens/sec 속도를 제시했다.
- 3.5 Flash Cyber는 보안 취약점 탐지용으로 설계되며 정부 및 신뢰 파트너를 대상으로 한 제한 파일럿 배포 계획이 있다.
토큰 사용량 감소와 경량 모델 옵션은 비용·지연을 동시에 낮춰 실서비스 적용 가능성을 높인다.
속도와 토큰 효율성 개선은 유의미하지만 실제 비용 절감·품질 향상 규모는 외부 벤치마크와 실사용 지표로 확인이 필요하다.
특히 Cyber처럼 제한 배포 모델은 접근성·검증성 문제로 인해 보안 적용 범위와 투명성에서 한계가 따른다.
원문 트윗 2개 보기
Google DeepMind
@GoogleDeepMind
We’re rolling out three new models to make AI agents faster, smarter, and cheaper at scale: Gemini 3.6 Flash: It uses fewer tokens than 3.5 Flash to deliver higher quality work at the exact same cost. Gemini 3.5 Flash-Lite: A fast, cost-effective option for everyday tasks like processing documents and agentic search. Gemini 3.5 Flash Cyber: A cybersecurity model built to find and patch critical software vulnerabilities.

Josh Woodward
@joshwoodward
Today’s launches are all about better performance, lower latency, and a smaller bill. + 3.6 Flash cuts token usage by up to 65% on complex coding + 3.5 Flash-Lite reaches speeds of 350 output tokens/sec Both are live in the Gemini app today! Next up: Gemini 3.5 Pro, which has officially entered partner testing.
We’re rolling out three new models to make AI agents faster, smarter, and cheaper at scale: Gemini 3.6 Flash: It uses fewer tokens than 3.5 Flash to deliver higher quality work at the exact same cost. Gemini 3.5 Flash-Lite: A fast, cost-effective option for everyday tasks
📈 NVIDIA Vera Rubin NVL72와 데이터센터·칩 레벨 효율성 주장포스트 5
NVIDIA는 Vera Rubin NVL72 플랫폼을 통해 메가와트당 토큰 처리량 개선을 전면에 내세웠고, CoreWeave와 DeepInfra의 초기 측정 결과를 인용해 기존 대비 높은 토큰/전력 효율과 동시성 개선을 보고했다. Spectrum‑6 스위치 등 네트워킹 요소도 생태계 전반의 처리량 향상 포인트로 제시됐다.
- CoreWeave의 첫 측정에서는 NVL72가 Blackwell 대비 메가와트당 토큰 처리량에서 10배 향상이 보고됐다.
- DeepInfra 벤치마크에서는 NVL72 기반 서버가 더 높은 처리량과 동시 에이전트 지원을 보였다고 보고됐다.
- 네트워킹 측면에서는 Spectrum‑6 스위치가 대규모 AI 팩토리의 병목 완화 요소로 소개됐다.
메가와트당 토큰 효율 개선은 대규모 운영 비용과 탄소 배출을 줄이는 직접적 수단이므로 데이터센터 운영자·서비스 사업자에게 의미가 크다.
측정치는 초기 실측 결과로 의미가 있으나 다양한 워크로드와 독립 벤치마크에서 재현될 필요가 있다.
원문 트윗 2개 보기
NVIDIA
@nvidia
The NVIDIA Vera Rubin platform is here, with 10x better performance per watt. The NVIDIA ecosystem, including @CoreWeave , @GoogleCloud , @Microsoft , and @Oracle Cloud, are standing up NVIDIA Vera Rubin NVL72 to deliver the lowest token cost for the agentic era. NVIDIA Vera Rubin NVL72 on CoreWeave demonstrates 10x more tokens per megawatt than Blackwell in their first measured performance. Benchmark results from @DeepInfra show that the NVIDIA Vera CPU is more than 2x as fast and can support more concurrent AI agents compared with other CPUs. Read more: https:// nvda.ws/4ywCiQn
NVIDIA
@nvidia
10x more tokens per megawatt. CoreWeave has the first measured performance of NVIDIA Vera Rubin NVL72, showing 10x improvement in tokens per second per megawatt on DeepSeek-R1 compared to Blackwell.
The first-ever measured silicon numbers for @NVIDIA Vera Rubin NVL72 are in First measured performance shows 10x more tokens per megawatt than Blackwell. No projections. Real results from live hardware.

📈 Poolside의 Laguna S 2.1 — 118B/8B MoE, 1M 토큰 컨텍스트, 오픈 가중치포스트 4
Poolside는 Laguna S 2.1을 118B 총파라미터·토큰당 8B 활성화 구조의 Mixture‑of‑Experts 모델로 공개했고, 최대 1M 토큰 컨텍스트와 오픈 가중치 배포를 통해 검증·수정 가능성을 강조했다. 자체 'Model Factory' 프로세스로 52일 만에 출시했다는 점도 주목점이다.
- 모델 스펙은 118B total, 8B activated per token의 MoE 구조이며 최대 1M 토큰 컨텍스트를 지원한다.
- 성능 지표로 TB 2.1 70%, SWE‑Bench Multi 78% 등 공개된 벤치마크 수치가 제시되었다.
- 가중치와 벤치마크 트래젝토리를 공개해 외부 검증과 재현을 유도했다.
오픈 가중치 공개는 연구·검증·커스터마이징을 가능하게 해 투명성과 재현성을 높인다.
대형 MoE 모델이 소형 모델 대비 성능을 넘어서 보이는 경우가 있어 추가 독립 검증과 재현성이 요구된다.
원문 트윗 2개 보기
Poolside
@poolsideai
Today we're releasing Laguna S 2.1, our most capable model to date. It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes. Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark. Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface https:// poolside.ai/blog/introduci ng-laguna-s-2-1 …
Jason Warner
@jasoncwarner
Today we’re releasing Poolside Laguna S 2.1 It is a 118B-total, 8B-active open-weight model built for agentic coding and long-horizon work, with context up to 1M tokens https:// poolside.ai/blog/introduci ng-laguna-s-2-1 … Laguna S 2.1 sits at the top of its weight class and competes with open models many times larger, while remaining small enough to run on a single NVIDIA DGX Spark It is available today under OpenMDW-1.1 via Hugging Face, OpenRouter, the Poolside API, and pool This is a remarkable model by any measure. As much as we can gather, it's the best open weight model in the West, regardless of size But the model itself is only part of the story There is a prevailing narrative that building capable models requires ever more capital, compute, and people. We believe the more important question is how efficiently you can turn those resources into intelligence At Poolside, we approach model building as an industrialized process spanning data, pre-training, reinforcement learning, evaluation, and inference. We call that system the Model Factory. We have written about our approach to model building extensively in a 6 part blog series: https:// poolside.ai/blog/introduci ng-the-model-factory … Laguna S 2.1 went from the Model Factory kicking it off to release in 52 days. The model is the output. The ability to keep building better models, faster and more efficiently each time is the actual innovation. The Model Factory is Poolside's compounding asset And so here we are, 52 days after kicking this off, releasing a 118B/8B MOE that tops 70% on TB 2.1, 78% on SWE-Bench Multi, 59% on SWE-Bench Pro, and 40% on DeepSWE. And it's fully open, from an American company I grew up in the Linux and Python communities, then spent much of my career at Canonical (Ubuntu), Heroku and GitHub. Those communities shaped my belief that important technology becomes more useful when people can understand it, challenge it, and build upon it. And even more important than being useful is being trusted. That is what open weights mean to me. They are not a marketing or distribution exercise. They give people control: the ability to inspect the work, reproduce the claims, modify the model, and run it inside their own environment I believe the West needs a credible open path to frontier intelligence. I believe it should be from an American company. We intend to be that company Laguna S 2.1 is another step in that direction and yet more evidence of Poolside's long held, often times contrarian beliefs and views on how intelligence will be built. One last note. We are doing something very different that we hope becomes industry norm going forward. We recognize that releasing a 118B/8B MOE that performs as well as Laguna S does would be met with some degree of skepticism. Models of this size are not supposed to outperform models 4-25x larger. We double and triple checked our benchmark trajectories to be sure. But we wanted to go another step and release those trajectories for you to see and help us quadruple check them. https:// trajectories.poolside.ai If you find something we missed, we genuinely want to hear about Have fun building whatever thing you can think of with the most persistent little model that could
📈 Tesla Robotaxi 영역 확장포스트 2
Tesla Robotaxi 서비스가 Orlando, Miami, Tampa, Dallas, Houston 등 7개 지역으로 확장되었고, 일부 지역에서는 완전 비감독(unsupervised) Model Y 운행이 이뤄지고 있다. Austin은 비감독·안전 모니터 혼합 운행, Bay Area는 안전 모니터 운행 위주로 보고되었다.
- 보고된 지역 목록에 Orlando·Miami·Tampa·Dallas·Houston·Austin·Bay Area가 포함된다.
- 운행 형태는 지역별로 비감독 Model Y 전용 또는 안전 모니터 병행 형태로 다양하다.
➖ Claude Cowork의 'Record a skill' 기능 도입포스트 1
Claude Cowork는 사용자가 화면 녹화와 함께 작업 과정을 설명하면 그 과정(스크립트)을 재실행 가능한 'skill'로 변환하는 기능을 추가했고, 데스크톱 앱의 '+ 메뉴'에서 'Record a skill'로 접근 가능하다. 기능은 Pro·Max·Team 플랜에서 사용할 수 있다.
- 기능은 화면 녹화와 음성 해설을 결합해 작업 시퀀스를 자동으로 캡처한다.
- 캡처된 시퀀스는 이후 Claude가 동일 작업을 수행하도록 재사용 가능한 skill로 변환된다.
- 배포 범위는 Pro, Max, Team 플랜으로 명시되었다.
📈 SkyPilot의 AI 컴퓨트 플랫폼 출시와 자금 조달포스트 1
SkyPilot은 분산된 GPU 자원 관리를 통합하는 AI 컴퓨트 플랫폼을 공개하며, 전용 고객 사례와 함께 2천만 달러 이상 규모의 투자 유치를 발표했다. 플랫폼은 페더레이션된 클러스터, 프리트레이닝·포스트트레이닝·서빙 워크로드 최적화 및 오픈소스 사용자의 플랫폼 전환 용이성을 강조한다.
- 목표는 컴퓨트 파편화(compute fragmentation)를 해소하고 GPU 활용률을 높이는 플랫폼 제공이다.
- 오픈소스 SkyPilot 사용자는 서버 URL 변경으로 플랫폼 전환이 가능하다고 명시되었다.
- 회사 발표에는 대규모 GPU 관리(예: 10,000+ GPUs)와 사용량 증가 수치가 포함되었다.
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.
