TL;DR
이번 포스트 묶음에서는 Ox Alpha와 Grok 4.6을 중심으로 코딩 모델의 성능·비용 경쟁과 무료 접근 확대가 이어졌고, Starlink·Terafab·궤도 데이터센터가 연결된 우주 인프라 투자 흐름도 나타났다. 에이전트 측면에서는 Claude Mythos 5 보안 스캔, NVIDIA의 실행 권한 분리, AVO의 장기 작업 평가가 모델 자체보다 운영 구조와 검증 절차를 중시하는 방향을 보였다. Physical AI에는 EgoSuite-Open100K처럼 상업 학습을 허용한 대규모 1인칭 데이터가 공급됐으며, AI 애플리케이션 개발에는 데이터 접지·에이전트 구축·평가 주도 개발·프로덕션 운영이 핵심 역량으로 묶였다. 모델 라우팅 비용, 사전 학습 데이터 반복, 저랭크 구조처럼 추론과 학습의 효율을 좌우하는 연구도 함께 부상했다.
𝕏 실시간 트렌드 토픽
🔥 Ox Alpha·Grok 4.6 코딩 모델의 성능과 무료 접근 경쟁포스트 9
Ox Alpha의 무료 배포와 Grok 4.6의 CursorBench 성적이 맞물리며 코딩 모델 경쟁이 점수뿐 아니라 태스크 비용과 사용량 제한까지 확장됐다.
- Ox Alpha는 OpenRouter, OpenCode, Cline에서 무료 사용 경로가 열렸고, 일부 포스트는 1M context·멀티모달 입력·zero data retention·near-unlimited usage를 내세웠다. 반면 Ox Alpha의 정체와 성능 수치는 독립 검증이 아니라 초기 벤치마크와 관측에 기반한 주장으로 남아 있다.
- Grok 4.6 extra high thinking mode는 CursorBench 3.2에서 70.8%, 태스크당 2.81달러로 인용됐으며 Fable 5 Max는 70.5%·17.32달러, Opus 5 Max는 70.0%·8.23달러, GPT-5.6 Sol Max는 67.2%·5.69달러로 비교됐다. 이 비교는 높은 점수와 낮은 비용을 함께 평가하는 구도다.
- Grok Build는 subagents·workflows·동시 멀티태스킹을 중심으로 거의 매일 개선되고 있으며, Grok Bot은 SuperGrok Plus·Cursor Pro+·Cursor Teams 구독자와 제한적 무료 체험 사용자에게 접근 범위를 넓혔다. 모델 경쟁이 API 성능만이 아니라 에이전트 작업 흐름과 배포 채널 경쟁으로 이동한 셈이다.
Ox Alpha 관련 포스트들은 무료 사용, 1M context, 멀티모달 처리와 코딩 벤치마크 성적을 근거로 단기간에 활용할 가치가 있는 모델이라고 평가했다.
Ox Alpha의 정체와 일부 벤치마크 수치는 확인되지 않은 초기 정보이며, Grok 4.6의 CursorBench 수치도 인용 포스트 기반이므로 동일 조건의 독립 평가가 더 필요하다.
원문 트윗 2개 보기

Elon Musk
@elonmusk
Grok 4.6 on extra high thinking mode now achieves #1 score on CursorBench!
Grok 4.6 just took the #1 spot on CursorBench 3.2.....and the efficiency is insane Here's the cost comparison: • Grok 4.6 Extra High — 70.8% | $2.81/task • Fable 5 Max — 70.5% | $17.32/task • Opus 5 Max — 70.0% | $8.23/task • GPT-5.6 Sol Max — 67.2% | $5.69/task Grok
Md Ismail Šojal
@0x0SojalSec
Ox Alpha free model beating Fable-5, GPT-5.6-sol, Grok 4.6, and GLM-5.3 on coding tasks. - Ox Alpha dropped 80% on a hard coding suite a 10-task DeepSWE subset. - Most frontier models are sitting at 52–65%. - 1M context, Multimodal, Zero data retention & Near-unlimited usage. Fingerprinting points strongly toward a next-gen GLM, but nothing is confirmed. This might be the strongest free coding model available right now. It’s free for the next week on OpenCode & OpenRouter.
A new frontier model just went fully free for a week. Claimed capacity: 100T tokens/day Ox Alpha just dropped as a free stealth model for 7 days: - 1M context window - (Multimodal) Text,Image,Video - Zero data retention - Near-unlimited usage Built for coding and x.com/opencode/statu…
📈 Starlink·Terafab·궤도 데이터센터로 이어지는 우주 AI 인프라포스트 5
Starlink 위성망 확장과 항공사 도입, Texas 반도체 제조 투자, 궤도 데이터센터 자금 조달이 하나의 인프라 축으로 묶였다.
- Elon Musk는 Starlink가 궤도에 1만 1,000개의 위성을 보유했다고 밝혔고, SpaceX는 Florida에서 Falcon 9로 Starlink 위성 29기를 발사했다. 위성 수 확대는 지상 인터넷 서비스와 발사 운영이 함께 커지는 구조로 나타났다.
- 항공사들의 Starlink 계약 발표가 늘면서 기내 연결성이 승객 선택에 영향을 줄 수 있다는 포스트가 나왔다. 원문은 Starlink를 제공하지 않는 항공사가 경쟁사로 승객을 잃을 위험을 이유로 들었지만, 구체적인 계약 항공사 수나 시장 점유율은 제시하지 않았다.
- Terafab 발표 뒤 Texas Brazos Valley에 추가 기업 문의가 들어왔고, 지역 지도자들은 향후 10년 지역 성장의 약 15%를 Terafab이 차지할 것으로 추산했다. Texas A&M의 microelectronics·semiconductors 석사 과정에는 올해 최소 50명이 시작할 전망이며, Starcloud는 2억 5,000만 달러를 조달해 23억 달러 기업가치를 기록하고 SpaceX Starship으로 궤도 데이터센터를 운용할 계획이다.
원문 트윗 2개 보기

Elon Musk
@elonmusk
Starlink now has 11k satellites in orbit
Watch Falcon 9 launch 29 @Starlink satellites to orbit from Florida https:// x.com/i/broadcasts/1 oKMvNQVqmkGQ …
TLDR Newsletter
@tldrnewsletter
Starcloud, a startup building data centers that run AI workloads in orbit, raised $250 million at a $2.3 billion valuation, up from $1.1 billion in March. NVIDIA and Cisco Investments joined the round, and Starcloud plans to fly its largest orbital data center spacecraft on SpaceX's Starship.
➖ Claude Mythos 5와 에이전트 실행 권한의 보안 설계포스트 5
에이전트 보안이 모델의 탐지 능력만이 아니라 실행 루프를 통제하는 harness와 실제 권한을 제한하는 infrastructure의 분리 문제로 확장됐다.
- Claude Security scans는 Claude Mythos 5에서 실행되며 Claude Enterprise 고객에게 public beta로 제공된다. 모델에 직접 접근하지 않아도 코드베이스를 스캔하고 결과를 반환하는 구조라서 보안 모델의 기능을 기존 제품 흐름 안에 연결한다.
- Claude Code on the web은 기존 팀 모델을 사용해 제안 패치를 열고, Mythos 5는 스캔 뒤 findings만 반환한다. 스캔 비용은 별도 모델 이용료가 아니라 기존 플랜의 표준 token usage로 청구되는 방식이다.
- NVIDIA 관련 포스트는 harness가 에이전트가 시도할 행동을 정하고 infrastructure가 실제로 가능한 행동을 통제한다고 구분했다. 이 구조는 에이전트가 계획을 세우는 계층과 파일·네트워크·배포 권한을 집행하는 계층을 나누는 방식이며, Claude는 Mythos 5 통합 파트너 확대와 오픈소스 보안용 3,500만 달러 크레딧 계획도 함께 밝혔다.
보안 스캔을 기존 Claude Enterprise와 Claude Code 흐름에 연결하고 findings만 반환하면 방어팀이 별도 모델 접근 없이 결과를 활용할 수 있다는 입장이다.
에이전트 안전성은 모델의 판단만으로 결정되지 않고 harness와 infrastructure가 각각 행동과 실제 권한을 통제해야 한다는 구조적 관점이다.
원문 트윗 2개 보기

Claude
@claudeai
Claude Security scans now run on Claude Mythos 5, available today in public beta for all Claude Enterprise customers. Put our most capable security model to work on your codebase, no separate model access needed.

NVIDIA AI
@NVIDIAAI
As AI agents take on more complex work, where should security live? The harness guides what an agent tries to do. Infrastructure controls what it can actually do. Our AI safety and security teams share how they’re thinking about security across the agent stack.
📈 EgoSuite-Open100K와 Physical AI 데이터 확장포스트 1
LightwheelAI가 1인칭 인간 행동 데이터를 대규모로 공개하면서 Physical AI 학습 데이터의 규모, 장면 다양성, 상업 이용 조건이 함께 부각됐다.
- EgoSuite-Open100K는 완전 주석 처리된 1인칭 인간 데이터셋으로, 총 10만 시간의 데이터와 1만 5,000개 이상의 작업·실제 장면을 포함한다. 공장 바닥부터 소매점 후방 공간까지 다양한 환경을 담고 손·몸 자세와 subtask-level semantics를 기록했다.
- 데이터는 일부 구간에 손목 카메라 시점을 포함하며 상업적 학습에 사용할 수 있는 라이선스로 공개됐다. 첫 1만 시간이 먼저 제공되고 나머지는 단계적으로 공개되는 입력 구조다.
- 원문은 사람 데이터가 Physical AI의 scaling law를 뒷받침하는 핵심 입력이라고 평가하며, 공개 데이터 위에 공동 foundation을 구축해야 한다고 밝혔다. 구체적인 학습 모델 성능이나 benchmark 점수는 제시되지 않았다.
➖ NVIDIA AVO와 장기 작업 에이전트의 평가 경계포스트 2
NVIDIA AVO가 기억·도구·실행 피드백을 이용한 장기 작업 구조로 소개됐지만, ARC-AGI-3 공개 시연 점수와 전체 benchmark 성적을 구분해야 한다는 지적이 함께 나왔다.
- AVO는 작업 중 지속적으로 상황을 검사하고 계획·실행·평가를 반복하며 memory, tools, execution feedback을 다음 단계에 반영한다. 매번 새 context에서 시작하지 않고 누적된 상태를 이용하는 방식이 장기 작업 지속성을 겨냥한다.
- 한 뉴스레터 포스트는 AVO가 ARC-AGI-3 공개 benchmark에서 183개 level과 25개 environment를 모두 통과해 100점을 기록했다고 전했다. ARC-AGI-3는 지시나 명시된 목표 없이 turn-based environment에 에이전트를 투입해 장기 자율성을 측정하는 평가로 설명됐다.
- François Chollet은 공개 demonstration set에서 100%를 기록한 것과 ARC-AGI-3 benchmark 전체에서 100%를 기록한 것은 같지 않다고 지적했다. 따라서 공개 시연 범위, 전체 평가 범위, 일반화 조건을 분리해 점수를 읽어야 한다.
AVO는 memory·tools·execution feedback을 반복 루프에 넣어 긴 작업에서 이전 상태를 유지하는 구조를 갖췄다는 평가다.
공개 demonstration set의 완벽한 점수를 전체 ARC-AGI-3 benchmark 성적처럼 해석하면 안 되며, 평가 범위와 일반화 조건을 구분해야 한다는 반론이다.
원문 트윗 2개 보기

François Chollet
@fchollet
This is very nice work from NVIDIA. Like all high-performing approaches on ARC-AGI-3, it uses deep learning-guided on-the-fly synthesis of symbolic world models, i.e. navigating the world by generating programs to represent what you know. To be clear, like with several other recent claims, scoring 100% on the public demonstration set is not the same as "scoring 100% on the ARC-AGI-3 benchmark". It would be like saying you beat a videogame because you cleared the tutorial level.
NVIDIA AVO continuously inspects, plans, implements, and evaluates, using memory, tools, and execution feedback to build on what it learns along the way. This allows the system to sustain progress across long-running tasks rather than starting over with each model context. Read
TLDR Newsletter
@tldrnewsletter
NVIDIA said its AVO agent architecture reached a perfect 100 score on the ARC-AGI-3 public benchmark, clearing all 183 levels across 25 environments. ARC-AGI-3 tests long-horizon autonomy by dropping agents into turn-based environments with no instructions or stated goals.
➖ AI 애플리케이션 개발의 핵심 역량과 어려운 평가 설계포스트 4
AI Engineering Skills Map은 모델 기초부터 프로덕션 운영까지의 개발 역량을 묶었고, 별도 포스트들은 실제 작업을 측정하려면 평가 과제의 난도와 조합을 설계해야 한다고 짚었다.
- Andrew Ng의 목록은 LLM foundations, 데이터로 모델을 grounding하는 방법, agentic systems 구축, evaluation-driven development, production 운영, machine learning foundations로 구성된다. 전통적 소프트웨어와 달리 AI 시스템의 출력이 예측 불가능하다는 차이가 이 역량 구분의 배경이다.
- 평가 과제의 난도를 조정하는 방법으로 더 많은 데이터, 검색으로 찾아야 하는 정보, 모든 사례를 통과해야 하는 completeness, 두 능력을 결합한 cross-domain task가 거론됐다. 약한 모델과 강한 모델을 함께 실행해 항상 통과하는 과제를 걸러내고, perfect pass@k가 좋은 학습 신호가 아닐 수 있다는 기준도 제시됐다.
- Cross-domain task는 에이전트의 약점을 드러내기 쉽지만 인위적이고 toy-like해질 수 있으며, completeness는 과도한 확인 행동을 유발할 수 있다. 평가가 현실 작업을 포착하려면 과제의 정보 탐색·범위·능력 조합을 함께 조정해야 한다.
원문 트윗 2개 보기
DeepLearning.AI
@DeepLearningAI
Traditional software is predictable. AI is not. What fundamental skills do developers need in order to build and deploy AI applications? Here's Andrew Ng's list: LLM foundations Grounding models with data Building agentic systems Evaluation-driven development Operating in production Machine learning foundations This is the second installment of the AI Engineering Skills Map. Read more and subcribe: https:// hubs.la/Q04tRNDQ0 #AI #MachineLearning #AIEngineering

Viv
@Vtrivedy10
launching another update to this skill soon - open holy grail question if anyone wants to riff! what goes in “How-to-make-tasks-harder[.]md”? some options in a list below part of making a good eval is making sure it captures the real world. another part is making the actual task hard so there’s something to hill climb one way to calibrate “hard” is running a weaker model and smarter model, if everything passes all the time that’s not great -> perfect pass @ k isn’t a good learning signal a list of things that can make tasks harder, would love to compile more and hear other strategies - add more data (not super bullish on this as models are great brute forcers) - add information that needs to be discovered by search (ex: look across multiple tables to find missing information) - completeness (add many cases that all need to be passed, this works well. it can incentivize bad behavior though like over checking) - cross-domain tasks that combine 2 abilities. this is the best imo, agents are not good at this but it can look artificial and toyish humans are still useful here but would love to chat on thoughts
📈 모델 라우팅·사전 학습 반복·저랭크 구조의 효율 연구포스트 4
새 연구 포스트들은 어떤 모델을 호출할지 결정하는 비용, 데이터 반복 허용 범위, 사전 학습과 RL의 역할 분담을 함께 다뤘다.
- Pandora’s Router는 전문 모델별 가치 추정이 무료라는 기존 가정을 문제로 삼고, 후보를 검사하고 추정하는 비용까지 포함한 최적 탐색 문제로 model routing을 정식화했다. Gaussian signal model 아래에서 입력과 전문 모델별로 추가 추정 비용을 지불할 가치가 있는지 판단하며, multi-LLM·retrieval-augmented specialist·가변 inference-time reasoning benchmark에서 expensive estimator 호출을 줄이면서 exhaustive estimation과 맞먹는 품질을 목표로 한다.
- 사전 학습 데이터 반복 연구에 관한 포스트는 token-per-parameter를 고정했을 때 큰 모델이 더 많은 반복을 견딜 수 있다는 관찰을 전했다. 반복되지 않은 데이터를 남은 부분에 넣어 전체 조건을 유지한 설정이며, 고품질 데이터의 반복 가능성과 모델 규모·학습률의 관계가 쟁점이다.
- 저랭크 기하에 관한 포스트는 pre-training이 기본 구조를 맞추고 RL이 그 형태를 미세 조정한다고 설명한다. 관측에 큰 공백이 있어 optimizer가 올바른 구조를 알 수 없으면 Fine-tuning이 저랭크 단계의 오류를 고치지 못한다는 주장이다. Ling-3.0-flash-dspark는 NVIDIA Blackwell GPU 4장에서 batch 1 기준 1,120 tok/s, 평균 TPOT 0.78ms, 1,000개 요청 기준 accept length 9.95를 기록했다고 밝혔다.
원문 트윗 2개 보기
DAIR.AI
@dair_ai
Banger paper from Google DeepMind. It's on the very hot topic of model routing. Routers assume the value estimate is free. Working out which specialist handles a query best costs money too, and it is often the larger cost in the pipeline. This new work formalizes model routing as Pandora's Box, the classic problem of optimal search when inspection is expensive. Under a Gaussian signal model the policies come out in closed form, telling you per specialist and per input whether refining the estimate is worth the price. Across a multi-LLM benchmark, retrieval-augmented specialists, and LLMs with variable inference-time reasoning, Pandora's Router matches exhaustive estimation quality while calling the expensive estimator far less often. When competing estimates are noisy, value-of-information reasoning raises the strategic specialist's utility at everyone else's expense. Paper: https:// arxiv.org/abs/2608.20316 Track more trending AI papers in our academy: https:// academy.dair.ai
Ant Ling
@AntLingAGI
Today we are open sourcing Ling-3.0-flash-dspark, a DSpark draft model built specifically for Ling-3.0-flash. On 4 NVIDIA Blackwell GPUs at batch 1, it delivered 1,120 tok/s, 0.78 ms mean TPOT, and an accept length of 9.95 across 1,000 requests.
용어 해설
- CursorBench
- — 코딩 작업에서 모델의 문제 해결 성능과 작업 비용을 함께 비교하는 평가 기준이다. 이번 포스트에서는 CursorBench 3.2 점수와 태스크당 비용이 함께 인용됐다.
- 에이전트 역량(Agentic Capability)
- — 모델이 단순 응답을 넘어 목표를 해석하고 도구·환경과 상호작용하며 여러 단계를 수행하는 능력이다. Ox Alpha와 frontier 모델 경쟁의 핵심 기준으로 언급됐다.
- 물리 AI(Physical AI)
- — 현실 환경에서 사람이나 로봇의 행동과 장면을 이해하고 작업을 수행하는 AI를 뜻한다. EgoSuite-Open100K는 사람의 1인칭 데이터를 이 분야의 입력으로 삼는다.
- 장기 작업 자율성(Long-horizon Autonomy)
- — 에이전트가 짧은 한 번의 응답이 아니라 긴 작업 과정에서 상태와 기억을 유지하며 다음 행동을 이어가는 능력이다. ARC-AGI-3와 NVIDIA AVO 관련 포스트에서 평가 대상이 됐다.
- 모델 라우팅(Model Routing)
- — 입력별로 여러 전문 모델 가운데 어느 모델을 호출할지 선택하는 방식이다. Pandora’s Router는 후보 평가 자체에 드는 비용까지 포함해 탐색 순서를 정한다.
- 저랭크 기하(Low-rank Geometry)
- — 모델이 사전 학습 과정에서 습득하는 저차원 구조와 형태를 뜻한다. 원문의 주장은 이 구조가 잘못되면 RL이나 Fine-tuning만으로 보정하기 어렵다는 것이다.
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.