본문으로 건너뛰기
X (Twitter)조회 2

Gemini 3.6 Flash 출시·NVL72 전력 효율성·Laguna S 2.1 오픈 가중치 MOE 발표

Gemini의 토큰 효율화, NVIDIA의 메가와트당 토큰 증가, Poolside의 118B/8B MoE 공개가 성능·효율성 중심 경쟁을 촉발했다.

이 요약은 AI가 원문을 분석해 생성했습니다. 정확한 내용은 원문 기준으로 확인하세요.

TL;DR

이번 기간에는 성능·효율성 개선을 전면에 둔 발표들이 집중됐다. Google DeepMind는 Gemini 3.6 Flash와 3.5 Flash‑Lite/Cyber를 내놓아 토큰 사용량 감소와 3.5 Flash‑Lite의 약 350 output tokens/sec급 속도를 강조했고, NVIDIA는 CoreWeave·DeepInfra 측 측정 결과를 인용해 Vera Rubin NVL72가 메가와트당 토큰 처리량에서 큰 폭의 개선을 보였다고 보고했다. Poolside는 118B 총파라미터에 토큰당 8B만 활성화되는 MoE 구조의 Laguna S 2.1을 공개하며 최대 1M 토큰 컨텍스트와 공개 가중치를 제시했다. 이러한 흐름은 전력·토큰 비용을 낮추면서도 장기 컨텍스트와 에이전트형 워크플로 지원을 확대하는 방향으로 업계가 재편되는 신호로 해석되나, 일부 모델·지표는 제한적 배포 또는 외부 검증 필요성이 병존한다.

𝕏 실시간 트렌드 토픽

🔥 Gemini 3.6 Flash와 3.5 Flash‑Lite·Cyber의 제품군 확장포스트 12

Google DeepMind는 Gemini 3.6 Flash를 통해 이전 Flash 모델 대비 토큰 사용량을 줄이는 방향으로 개선을 가했고, 3.5 Flash‑Lite는 약 350 output tokens/sec급 속도와 저비용 옵션을 목표로 설계됐다. 3.5 Flash Cyber는 코드 보안 취약점 탐지에 특화되어 제한적 파트너 배포 방침이 적용됐다.

세부 내용 보기
  • 3.6 Flash는 3.5 Flash 대비 토큰 사용량을 줄이는 데 중점을 둔 업그레이드임이 강조됐다.
  • 3.5 Flash‑Lite는 빠른 응답이 필요한 에이전트형 워크플로를 겨냥해 약 350 output tokens/sec 속도를 제시했다.
  • 3.5 Flash Cyber는 보안 취약점 탐지용으로 설계되며 정부 및 신뢰 파트너를 대상으로 한 제한 파일럿 배포 계획이 있다.
찬성다수

토큰 사용량 감소와 경량 모델 옵션은 비용·지연을 동시에 낮춰 실서비스 적용 가능성을 높인다.

중립다수

속도와 토큰 효율성 개선은 유의미하지만 실제 비용 절감·품질 향상 규모는 외부 벤치마크와 실사용 지표로 확인이 필요하다.

반대소수

특히 Cyber처럼 제한 배포 모델은 접근성·검증성 문제로 인해 보안 적용 범위와 투명성에서 한계가 따른다.

원문 트윗 2개 보기

📈 NVIDIA Vera Rubin NVL72와 데이터센터·칩 레벨 효율성 주장포스트 5

NVIDIA는 Vera Rubin NVL72 플랫폼을 통해 메가와트당 토큰 처리량 개선을 전면에 내세웠고, CoreWeave와 DeepInfra의 초기 측정 결과를 인용해 기존 대비 높은 토큰/전력 효율과 동시성 개선을 보고했다. Spectrum‑6 스위치 등 네트워킹 요소도 생태계 전반의 처리량 향상 포인트로 제시됐다.

세부 내용 보기
  • CoreWeave의 첫 측정에서는 NVL72가 Blackwell 대비 메가와트당 토큰 처리량에서 10배 향상이 보고됐다.
  • DeepInfra 벤치마크에서는 NVL72 기반 서버가 더 높은 처리량과 동시 에이전트 지원을 보였다고 보고됐다.
  • 네트워킹 측면에서는 Spectrum‑6 스위치가 대규모 AI 팩토리의 병목 완화 요소로 소개됐다.
찬성다수

메가와트당 토큰 효율 개선은 대규모 운영 비용과 탄소 배출을 줄이는 직접적 수단이므로 데이터센터 운영자·서비스 사업자에게 의미가 크다.

중립다수

측정치는 초기 실측 결과로 의미가 있으나 다양한 워크로드와 독립 벤치마크에서 재현될 필요가 있다.

원문 트윗 2개 보기

📈 Poolside의 Laguna S 2.1 — 118B/8B MoE, 1M 토큰 컨텍스트, 오픈 가중치포스트 4

Poolside는 Laguna S 2.1을 118B 총파라미터·토큰당 8B 활성화 구조의 Mixture‑of‑Experts 모델로 공개했고, 최대 1M 토큰 컨텍스트와 오픈 가중치 배포를 통해 검증·수정 가능성을 강조했다. 자체 'Model Factory' 프로세스로 52일 만에 출시했다는 점도 주목점이다.

세부 내용 보기
  • 모델 스펙은 118B total, 8B activated per token의 MoE 구조이며 최대 1M 토큰 컨텍스트를 지원한다.
  • 성능 지표로 TB 2.1 70%, SWE‑Bench Multi 78% 등 공개된 벤치마크 수치가 제시되었다.
  • 가중치와 벤치마크 트래젝토리를 공개해 외부 검증과 재현을 유도했다.
찬성다수

오픈 가중치 공개는 연구·검증·커스터마이징을 가능하게 해 투명성과 재현성을 높인다.

반대소수

대형 MoE 모델이 소형 모델 대비 성능을 넘어서 보이는 경우가 있어 추가 독립 검증과 재현성이 요구된다.

원문 트윗 2개 보기

Poolside

@poolsideai

1달 전

Today we're releasing Laguna S 2.1, our most capable model to date. It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes. Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark. Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface https:// poolside.ai/blog/introduci ng-laguna-s-2-1 …

트윗에 첨부된 이미지
💬 4 30 52👁 2417

Jason Warner

@jasoncwarner

1달 전

Today we’re releasing Poolside Laguna S 2.1 It is a 118B-total, 8B-active open-weight model built for agentic coding and long-horizon work, with context up to 1M tokens https:// poolside.ai/blog/introduci ng-laguna-s-2-1 … Laguna S 2.1 sits at the top of its weight class and competes with open models many times larger, while remaining small enough to run on a single NVIDIA DGX Spark It is available today under OpenMDW-1.1 via Hugging Face, OpenRouter, the Poolside API, and pool This is a remarkable model by any measure. As much as we can gather, it's the best open weight model in the West, regardless of size But the model itself is only part of the story There is a prevailing narrative that building capable models requires ever more capital, compute, and people. We believe the more important question is how efficiently you can turn those resources into intelligence At Poolside, we approach model building as an industrialized process spanning data, pre-training, reinforcement learning, evaluation, and inference. We call that system the Model Factory. We have written about our approach to model building extensively in a 6 part blog series: https:// poolside.ai/blog/introduci ng-the-model-factory … Laguna S 2.1 went from the Model Factory kicking it off to release in 52 days. The model is the output. The ability to keep building better models, faster and more efficiently each time is the actual innovation. The Model Factory is Poolside's compounding asset And so here we are, 52 days after kicking this off, releasing a 118B/8B MOE that tops 70% on TB 2.1, 78% on SWE-Bench Multi, 59% on SWE-Bench Pro, and 40% on DeepSWE. And it's fully open, from an American company I grew up in the Linux and Python communities, then spent much of my career at Canonical (Ubuntu), Heroku and GitHub. Those communities shaped my belief that important technology becomes more useful when people can understand it, challenge it, and build upon it. And even more important than being useful is being trusted. That is what open weights mean to me. They are not a marketing or distribution exercise. They give people control: the ability to inspect the work, reproduce the claims, modify the model, and run it inside their own environment I believe the West needs a credible open path to frontier intelligence. I believe it should be from an American company. We intend to be that company Laguna S 2.1 is another step in that direction and yet more evidence of Poolside's long held, often times contrarian beliefs and views on how intelligence will be built. One last note. We are doing something very different that we hope becomes industry norm going forward. We recognize that releasing a 118B/8B MOE that performs as well as Laguna S does would be met with some degree of skepticism. Models of this size are not supposed to outperform models 4-25x larger. We double and triple checked our benchmark trajectories to be sure. But we wanted to go another step and release those trajectories for you to see and help us quadruple check them. https:// trajectories.poolside.ai If you find something we missed, we genuinely want to hear about Have fun building whatever thing you can think of with the most persistent little model that could

💬 7 13 51👁 8126

📈 Tesla Robotaxi 영역 확장포스트 2

Tesla Robotaxi 서비스가 Orlando, Miami, Tampa, Dallas, Houston 등 7개 지역으로 확장되었고, 일부 지역에서는 완전 비감독(unsupervised) Model Y 운행이 이뤄지고 있다. Austin은 비감독·안전 모니터 혼합 운행, Bay Area는 안전 모니터 운행 위주로 보고되었다.

세부 내용 보기

Claude Cowork의 'Record a skill' 기능 도입포스트 1

Claude Cowork는 사용자가 화면 녹화와 함께 작업 과정을 설명하면 그 과정(스크립트)을 재실행 가능한 'skill'로 변환하는 기능을 추가했고, 데스크톱 앱의 '+ 메뉴'에서 'Record a skill'로 접근 가능하다. 기능은 Pro·Max·Team 플랜에서 사용할 수 있다.

세부 내용 보기
  • 기능은 화면 녹화와 음성 해설을 결합해 작업 시퀀스를 자동으로 캡처한다.
  • 캡처된 시퀀스는 이후 Claude가 동일 작업을 수행하도록 재사용 가능한 skill로 변환된다.
  • 배포 범위는 Pro, Max, Team 플랜으로 명시되었다.
원문 트윗 1개 보기

📈 SkyPilot의 AI 컴퓨트 플랫폼 출시와 자금 조달포스트 1

SkyPilot은 분산된 GPU 자원 관리를 통합하는 AI 컴퓨트 플랫폼을 공개하며, 전용 고객 사례와 함께 2천만 달러 이상 규모의 투자 유치를 발표했다. 플랫폼은 페더레이션된 클러스터, 프리트레이닝·포스트트레이닝·서빙 워크로드 최적화 및 오픈소스 사용자의 플랫폼 전환 용이성을 강조한다.

세부 내용 보기
  • 목표는 컴퓨트 파편화(compute fragmentation)를 해소하고 GPU 활용률을 높이는 플랫폼 제공이다.
  • 오픈소스 SkyPilot 사용자는 서버 URL 변경으로 플랫폼 전환이 가능하다고 명시되었다.
  • 회사 발표에는 대규모 GPU 관리(예: 10,000+ GPUs)와 사용량 증가 수치가 포함되었다.
원문 트윗 1개 보기

Zongheng Yang

@zongheng_yang

1달 전

Today, SkyPilot is out of stealth. Building custom intelligence is now existential. We help frontier AI teams build intelligence faster by removing their biggest bottleneck: AI compute fragmentation. Frontier teams like @appliedcompute , @AbridgeHQ , @hippocraticai , @hcompany_ai , and @nubank already run on SkyPilot, with 10x faster time-to-intelligence and double-digit increase in GPU utilization. AI teams today get compute anywhere they can. They then firefight compute fragmentation across providers. Researchers burn time on workload setup. Infra gets paged when GPUs go down. Frontier teams build slowly even on the fastest compute. @skypilot_org turns your fragmented compute into one AI supercomputer, so you run frontier workloads faster. Many users manage 10,000+ GPUs across providers with SkyPilot. GPU hours consumption has grown 6x in the last 6 months. 1/ We're launching SkyPilot Platform — the AI compute platform for frontier AI teams to manage large GPU fleets and accelerate building custom intelligence. Optimized for fleet management, team governance, and frontier workloads — pretraining, post-training, multi-cluster serving, and sandboxes. SkyPilot open source users can switch to the platform with a server URL change. 2/ We've raised over $20M led by @Lux_Capital ( @breeves08 ), with participation from @AmplifyPartners ( @dauber , @lennypruss ), @coatuemgmt , @FoundationCap ( @ashugarg , @JayaGup10 ), @RaceCapital , @thehousefund , and top operators like @alighodsi (CEO, Databricks), @JeffDean (Chief Scientist, Google), @rauchg (CEO, Vercel), @amasad (CEO, Replit), @ClemDelangue (CEO, @huggingface ) and more. We're hiring across Engineering and GTM to deliver the platform for the next decade of AI. Above all, I'm excited to be building with the incredible team we've assembled, along with my cofounders Zhanghao @Michaelvll1 , Romil @bromil101 , Scott, and Ion @istoica05 . If you firefight AI compute, let's build.

💬 9 31 59👁 3495

용어 해설

Mixture-of-Experts (MoE)
MoE는 토큰별로 활성화되는 일부 전문가(expert) 서브네트워크만 연산에 참여시키는 아키텍처다. 전체 파라미터는 크지만 토큰당 활성화되는 파라미터가 작아 계산 효율과 모델 용량 확장이 가능하다. Laguna S 2.1처럼 '118B total, 8B activated' 구조를 구현할 때 핵심 개념이다.
컨텍스트 윈도우(Context Window)
컨텍스트 윈도우는 모델이 한 번에 처리할 수 있는 입력 토큰의 최대 길이다. 길이가 길수록 긴 문서·대화·코드의 문맥을 유지할 수 있으나 메모리·연산 비용이 증가한다. Laguna S 2.1의 최대 1M 토큰 컨텍스트는 장기 문맥 작업을 겨냥한 설계다.
메가와트당 토큰 처리량(Tokens per megawatt)
메가와트당 토큰 처리량은 제공 전력 대비 초당 처리 가능한 토큰 수를 뜻하는 에너지 효율 지표다. 데이터센터·칩 설계 관점에서 전력 효율성과 토큰 처리량의 균형을 평가할 때 사용된다. NVIDIA 측 발표는 NVL72의 메가와트당 토큰 효율이 대폭 개선되었다고 보고했다.
토큰 효율성(Token Efficiency)
토큰 효율성은 동일 작업을 수행할 때 소모하는 토큰 수와 관련된 성능 지표다. 토큰 사용을 줄이면 비용과 지연이 낮아지므로 모델 설계와 프롬프트 전략에서 중요한 최적화 대상이다. Gemini 3.6 Flash는 이전 세대보다 토큰 사용량을 줄이는 점을 핵심 개선점으로 삼았다.
AI 분석 전체 내용 보기

AI 요약 · 북마크 · 개인 피드 설정 — 무료

출처 · 인용 안내

원문 발행 2026. 07. 22.수집 2026. 07. 22.출처 타입 TWITTER

인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.