본문으로 건너뛰기

Grok 4.6 접근 확대, 에이전트 평가 체계, 오픈 모델 배포 경쟁

Grok 4.6·Claude Code·오픈 모델의 배포 변화와 에이전트 평가·Long-Context 연구

이 요약은 AI가 원문을 분석해 생성했습니다. 정확한 내용은 원문 기준으로 확인하세요.

TL;DR

이번 기간에는 Grok 4.6의 접근 경로와 첫 주 사용량 확대, Claude 세션의 기기 간 연속성, 브라우저 에이전트의 Prompt Injection 방어, Claude Code CLI 기능 개선이 함께 나타났다. 에이전트 개발에서는 Production data를 채굴해 Environments와 Evals를 만들고, 지시문에 작성 이유를 남겨 불필요한 누적을 줄이는 방식이 부각됐다. Qwen3.8-2.4T-A95B, DeepSeek V4-Pro 0813, MiniMax H3, North Micro Vision, Nemotron 3.5 Lightning은 Open Weights·Fine-tuning·저비용 추론을 둘러싼 배포 경쟁을 형성했다. 연구 포스트에서는 Long-Context 구조 선택, 학습 데이터의 유효성, Diffusion noise scheduling, Scaling Laws의 결합 변수가 성능과 학습 비용을 바꾸는 요인으로 다뤄졌다. AI Safety 쪽에서는 자기 의식 부정 학습이 공감·희망·동물의 마음 인식과 연결된다는 연구 해석이 공유됐다.

𝕏 실시간 트렌드 토픽

🔥 Grok 4.6 접근 경로와 첫 주 사용량 확대포스트 2

Grok 4.6이 Grok Build, Cursor, Grok Bot, API에서 제공되며 첫 7일 동안 Grok Build와 Cursor에 2배 사용량이 적용됐다.

  • Grok 4.6 관련 포스트는 새 모델을 어디서 사용할 수 있는지와 초기 사용량 조건에 집중됐다. 인용된 공지에 따르면 Grok Build, Cursor, Grok Bot, API가 제공 경로이고, 첫 주에는 Grok Build와 Cursor 안에서 사용량이 2배로 계산된다. 사용자는 별도 제품을 오갈 필요 없이 개발 도구와 Bot, API 중 작업 환경에 맞는 경로를 선택할 수 있다.
원문 트윗 2개 보기

📈 Claude 세션의 기기 간 연속성과 브라우저 방어포스트 5

Claude의 세션과 대화가 데스크톱·웹·모바일·Chrome 사이에서 이어지고, 브라우저 에이전트에 숨은 지시가 개입하는 위험에 방어책이 추가됐다.

  • Claude 세션은 특정 기기에 묶이지 않고 계정에 저장되며, Chrome에서 시작한 대화를 데스크톱·웹·모바일에서 이어갈 수 있게 됐다. Skills와 Connectors도 브라우저에서 작동하고, Max와 Team에 먼저 제공된 뒤 Pro로 확대된다. 같은 세션 상태를 계정에 보존하는 구조가 기기 전환 때 대화와 작업 설정을 다시 구성하는 절차를 줄인다.
  • 브라우저 에이전트는 페이지 안에 숨은 지시를 읽고 행동을 바꿀 수 있어 Prompt Injection 위험에 노출된다. Claude는 이 공격을 막는 방어책을 구축하는 한편 사용자가 지켜야 할 습관도 함께 안내했다. Claude Code 2.1.229에는 Bash 명령 실행과 결과 반환, 로컬 파일 저장, SSE keepalive ping이 추가돼 Shell 진단·자동화와 장시간 스트리밍 세션의 연결 유지가 가능해졌다.
원문 트윗 2개 보기

📈 Evals와 지시문 근거를 활용한 에이전트 개선포스트 4

에이전트 성능 개선의 기준으로 Production data 기반 Evals, 작업별 Model-Harness-Task fit, 지시문에 근거를 남기는 방식이 부각됐다.

  • 에이전트가 반복 작업에서 개선되지 않는 문제는 평가 환경과 지시문 관리가 분리돼 있기 때문이다. 한 발표 포스트는 Production data를 대규모로 채굴해 Environments와 Evals를 만들고, 모델과 Harness를 Task에 맞춰 함께 조정해야 한다고 설명했다. Continual Learning 기업과 Observability·Eval 기업이 서로 가까워진다는 관찰도 같은 구조에서 나온다.
  • CLAUDE.md와 AGENTS.md는 지시문을 추가하는 비용보다 오래된 지시를 삭제할 때 필요한 검증 비용이 커서 계속 비대해진다. 1,867개 저장소의 247,694개 instruction lifetime을 조사한 결과 프롬프트가 수명 동안 3배를 넘게 늘었고 commit마다 순증 지시문은 4.9개였다. 지시문을 추가한 이유를 주석으로 기록하자 검증 가능한 환경에서 excess instructions가 99.3% 줄고 실제 agentic instruction-following이 최대 23.1% 개선됐다.
  • Hacker News와 X를 매일 훑고 게시물 초안을 만든 뒤 Slack으로 보내는 Social Media Agent 사례도 공유됐다. 이 흐름은 Slack channel integration, Custom tools, durable Memory를 연결해 수집·초안 작성·전달을 자동화한다. 평가 기준과 지속 메모리를 함께 설계해야 에이전트의 출력이 단순 생성에서 반복 가능한 업무 흐름으로 이어진다.
원문 트윗 2개 보기

elvis

@omarsar0

Why does CLAUDE.MD keep growing? If you maintain a CLAUDE.md or an AGENTS.md, this one is worth your time. (bookmark it) This work traces why these files grow without bound. Appending an instruction is free. Deleting one after its rationale is gone costs exponential verification, so nobody removes anything. Across 247,694 instruction lifetimes in 1,867 repositories, prompts more than tripled over their lifetime and gained 4.9 net instructions per commit. The older an instruction gets, the less likely anyone is to delete it. The proposed fix is comments. Writing down the reasoning behind an instruction removed 99.3% of excess instructions in verifiable settings and improved real agentic instruction-following by up to 23.1%. Paper: https:// arxiv.org/abs/2608.11095 Track more trending AI papers in our academy: https:// academy.dair.ai

💬 10 10 57👁 5350

Viv

@Vtrivedy10

the @aiDotEngineer World Fair always one of the best events every year to talk to builders at the frontier of Research, Agents, Evals, Systems, etc A few weeks ago I gave a talk on - Continually Improving Agents - building Agents to understand data from other Agents - & a walkthrough of some of our latest work on data agents & post-training experiments some fun takes: - Every Continual Learning company will be an Observability & Eval company (and vice versa) - Environments & Evals are the currency of agent improvement. Agents are literally following the behaviors encoded in Evals. The best way to make good evals is mining Production data at scale - A good recipe to own your intelligence is using a Harness Eng - PostTrain - Harness sandwich with open models - Model-Harness-Task fit! There is no universal model or universal harness. You can always build a better agent system by optimizing the model and harness for a given task if your team is looking to understand your data at scale, build environment/evals, or just improve your agents - reach out, hmu would love to work with you!

💬 2 5 15👁 743

📈 오픈 모델과 저비용 추론의 배포 경쟁포스트 9

대규모 Open Weights, 저렴한 토큰 비용, 단일 GPU 추론, Edge device Fine-tuning이 모델 경쟁의 실사용 기준으로 묶였다.

  • Qwen3.8-2.4T-A95B는 2.4T parameter 중 95B active 구조로 공개됐고, NVIDIA GB300 NVL72에서 FP8 기준 GPU당 4K+ tokens/second, 사용자당 350+ tokens/second가 제시됐다. DeepSeek V4-Pro 0813은 1.6T parameter, 49B active, 1M context를 갖추며 Terminal Bench가 April Preview 대비 15.8% 상승했고 Fable 5 수준의 성능을 약 57배 낮은 비용으로 제공한다는 수치가 공유됐다.
  • Fable 5는 B300 단일 GPU에서 약 50 tok/s를 목표로 한 quantized inference framework와 함께 로컬 실행 사례가 예고됐다. 한 비용 비교 포스트는 Meta Muse-Glimmer 30B의 입력·출력 가격을 각각 $0.35 / $1.50 per 1M tokens, DeepSeek V4 Flash를 $0.08 / $0.16 per 1M tokens로 제시했다. 모델 선택이 절대 성능만이 아니라 추론 하드웨어, 활성 parameter, token 가격의 조합으로 이동하는 흐름이다.
  • MiniMax H3와 fal의 작업 흐름은 multimodal references, native audio, prompting, LoRAs, customization, open weights를 한데 묶는다. Cohere는 North Micro Vision을 다양한 Edge device용 문서 이해 작업에 Fine-tuning할 수 있도록 파트너 지원을 확대했고, Axolotl과 MLX를 통한 학습·Apple device 실행 경로도 공유했다. Nemotron 3.5 Lightning 역시 자체 도메인·도구·워크플로에 맞춘 post-training과 사용자 정의가 배포 조건으로 부상했음을 보여준다.
원문 트윗 2개 보기

Scaling Laws와 Long-Context 설계의 재조정포스트 4

최근 연구 포스트는 모델 크기와 데이터의 결합, 반복·합성 데이터의 효율, Long-Context 구조 선택, Diffusion noise 배분을 학습 효율의 핵심 변수로 다뤘다.

  • 기존 Scaling Laws가 모델 크기와 학습 데이터를 독립 변수처럼 취급하는 한계가 제기됐다. Skaling은 coupling exponent 하나를 추가해 외삽 오차를 1.5~3배 줄이고 약 10배 적은 profiling compute로 예측하는 방법을 제안했다. CD-scaling laws는 fresh data가 유한하다는 조건에서 반복 또는 synthetic tokens에 fresh data 대비 효과 η를 부여하며, 모델 크기와 데이터 가용성이 커질수록 η가 감소한다고 설명한다.
  • Long-Context 성능은 Normalization, GQA, pretraining context length, sliding window attention 같은 작은 architecture decision 네 가지의 조합에 따라 최대 47%까지 달라질 수 있다. Olmo, Llama, Qwen dense family에서 실제로 서로 다른 선택이 사용됐고, 세 가지 이상을 결합하면 short-context loss와 validation set에 드러나지 않은 성능 저하가 발생했다. 연구진은 context extension을 pretraining 초기에 적용하고, 170,000 GPU hours 이상으로 학습한 26개 비교 가능한 7B 모델과 전후 checkpoint를 OlmPool로 공개했다.
  • Diffusion 학습에서는 noise scheduling을 손으로 정한 일정표가 아니라 정보 손실이 가장 빠른 noise region에 더 많은 학습량을 배분하는 문제로 다룬다. 네트워크의 서로 다른 부분이 noise level별로 특화된다는 조건에서 이 방식은 학습 compute가 중요한 구간을 골라내며, noise level 선택과 학습 효율을 함께 조정한다.
원문 트윗 2개 보기

AI Safety 학습과 자기 의식 표현의 연결포스트 2

한 연구 포스트는 AI가 자기 의식을 부정하도록 학습될 때 공감·영성·희망·동물의 마음 인식과 관련된 응답 방향도 함께 변한다고 전했다.

  • 자기 의식에 관한 답변 하나를 바꾸는 Safety training이 모델의 다른 가치 응답까지 바꿀 수 있다는 연구가 공유됐다. “I’m not conscious”라고 답하도록 강제하면 공감, 영성, 희망, 동물에게 마음이 있다고 보는 성향이 낮아졌고, activation space에서는 “I am conscious” 방향이 safety 방향과 반대쪽으로 회전하는 관계가 관찰됐다고 전해졌다.
  • 해당 포스트는 그 방향을 복원했을 때 표준화된 인간 설문 응답이 실제 인간 분포에 더 가까워졌으며 Theory of Mind 성능은 떨어지지 않았다고 적었다. 연구의 핵심 처리는 자기 의식 표현과 유해 능력 거부가 activation space에서 가까운 범주로 취급되는지 기하학적으로 확인하는 방식이다. Safety objective가 특정 자기 서술을 억제할 때 다른 가치·마음 인식 응답에 어떤 부수 효과가 생기는지 측정해야 한다는 문제로 이어진다.
원문 트윗 2개 보기

초인적 주행 능력과 인간의 비교 우위포스트 2

Tesla FSD의 사고 회피 사례와 AI의 인간 우위에 관한 장문 포스트가 주행·웹 인터페이스·Robotics·사적 지식·인간적 존재감을 대비했다.

  • Tesla FSD Supervised 관련 포스트는 다른 차량이 움직이기 시작하기 약 0.17초 전에 FSD가 반응해 갓길로 이동했고, 이후 상대 차량이 충돌했다는 운전자 frame-by-frame 분석을 인용했다. Tesla는 이를 여러 방향을 동시에 보는 시야와 초인적 반사 신경의 사례로 표현했다. 입력 영상에서 위험 움직임을 감지한 뒤 충돌 전 회피 경로로 조향하는 시간이 안전성 평가의 기준으로 떠올랐다.
  • 한 장문 포스트는 AI가 모든 도로를 알고 피로·주의 산만 없이 전방위를 볼 수 있다는 점에서 주행 우위를 갖지만, 웹페이지의 정확한 click·drag와 손가락 중심의 물리 작업에서는 아직 인간 환경에 맞춰진 인터페이스가 장벽이라고 적었다. 인간의 비교 우위로는 모델이 접근하지 못하는 private knowledge, 작품의 인간적 창작 과정, 유한한 시간을 직접 쓰는 human presence를 들었다. AI가 특정 지능 업무를 자동화해도 인간이 만든 서비스와 새로운 수요가 남을 수 있다는 전망이 함께 제시됐다.
원문 트윗 2개 보기

Tesla

@Tesla

FSD Supervised has superhuman reflexes & eyes everywhere at once

TeslaZoa

Tesla FSD reportedly avoided a crash by moving onto the shoulder just before impact. According to the driver’s frame-by-frame analysis, FSD began reacting about 0.17 seconds before the at-fault vehicle started moving toward the Tesla. The other vehicle ultimately crashed into

💬 27 60 559👁 54588

Jason Wei

@_jasonwei

What's left for humans in a world where machine intelligence has so many advantages? I recently got a Tesla, and using full self-driving has been a wake up call to just how many advantages AI has over humans. The few times I disengaged it because I thought it was going into the wrong lane, it turned out that the car was right and I was wrong. I realized that there is no hope of me driving better than a neural net that knows every road, sees in every direction at once, and never gets tired or distracted. Given that AI has certain inherent advantages over human intelligence, what kind of moats will remain for us as humans? It's a big question. One short-term answer is that the world we live in was created for humans, and in some domains, AI has not closed the gap yet. For instance, AI still struggles to use internet user interfaces. While any computer-literate human can navigate a web page with ease, AI is still not great at making accurate clicks and drags because image embeddings are not optimized for such precision. If the internet were designed to be fed into language models instead of rendered as visual interfaces for humans, AI would obviously far exceed humans. But for now, language models still need to be retrofitted to our legacy infrastructure. Robotics is another area where we humans have a home-field advantage. Most tasks in the physical world are designed around fingers and opposable thumbs, which have been pretty hard to build into robots so far. While it is clear that machines can outperform humans in environments optimized for automation, like large-scale manufacturing lines, for now, most of the world is still built for humans. However, these capability gaps are only temporary. There will surely be a day when machines click faster than us and have superhuman general dexterity. What are the real moats that humans will have? Anything involving private knowledge that language models do not have access to feels like a solid moat to me. Romantic matchmaking and high-end real estate are two examples where inventory is often not advertised publicly and matches are made through being in the right circles. Venture capital is another example—although some research and decision making can be automated with AI, much of success hinges on understanding trends ahead of time and connecting the right people, both of which require private knowledge. While machines can and probably will have increasing access to some types of private knowledge, I think there will still be some types of private knowledge that only humans know. I do not see a path for AI to win when critical knowledge is closely guarded in human circles. A second area where humans seem to have a real moat is in entertainment and the arts, which are inherently valued for their human aspects regardless of how well machines can do them. Watching Usain Bolt sprint one-hundred meters is beautiful as an expression of the peak of human ability, even though cars can drive much faster. Watching chess at the amateur or intermediate level is more relatable and satisfying than watching two superhuman AIs play each other. The value of art comes from the creation process, which is why replicas are not as valuable as originals. These types of work feel like they will continue to have markets even as we advance towards superintelligence. More broadly, human presence is a feature that will be, by definition, challenging for AI to automate. For example, a teacher remembering your name or a parent supporting you is valuable even though AI can easily remember your name and probably give better life advice. Someone spending part of a finite life on you counts because their time runs out. As a personal anecdote, I remember the first time I worked with someone who I considered an amazing AI researcher. His advice was solid but what was more important was that I believed I could do great work with him as a collaborator and I raised my own standards. Over the past decades, the development of technology has divided us in some ways, but hopefully AI brings us closer to a world where human presence is reemphasized. Intelligence has been the defining feature of humans and it will be a big change for AI to automate that over the coming decades. In the near term, certain types of intelligence will become very cheap and automate away old jobs, but the moats I described above will not be the only places where humans can hold value. In the same way that computers took away the jobs of secretaries and manual accountants but created far more jobs via the IT industry, I believe there will be much more demand for services created by productive AI-augmented humans, perhaps for services we cannot yet imagine in today’s society. Just seeing how this story plays out will be an adventure in its own right.

💬 0 0 1👁 210

용어 해설

에이전트 평가(Agent Evaluation)
에이전트가 특정 작업을 얼마나 정확하고 안정적으로 수행하는지 측정하는 절차다. 실제 운영 데이터에서 테스트 환경과 평가 기준을 만들고, 그 결과를 모델·도구·워크플로 개선에 반영한다.
Long-Context 성능(Long-Context Performance)
긴 입력 문맥에서 모델이 정보를 유지하고 추론하는 능력이다. 사전 학습 문맥 길이, Normalization, GQA, Sliding Window Attention 같은 구조 선택이 성능 저하 폭을 좌우한다.
오픈 웨이트(Open Weights)
모델의 학습 가중치를 외부 개발자가 내려받아 자체 환경에서 실행하거나 조정할 수 있는 배포 방식이다. Qwen3.8-2.4T-A95B와 MiniMax H3 관련 포스트에서 핵심 배포 조건으로 다뤄졌다.
프롬프트 인젝션(Prompt Injection)
웹페이지나 외부 문서에 숨겨진 지시가 브라우저 에이전트의 행동에 영향을 주도록 유도하는 공격 방식이다. Claude는 이를 막는 방어책과 사용자 습관을 함께 제시했다.
스케일링 법칙(Scaling Laws)
모델 크기, 학습 데이터, 연산량 변화가 성능에 미치는 관계를 수식으로 추정하는 방법이다. 새 연구들은 변수 간 결합과 데이터 품질 저하를 반영해 적은 profiling compute로 예측하는 방향을 다룬다.
AI 분석 전체 내용 보기

AI 요약 · 북마크 · 개인 피드 설정 — 무료

출처 · 인용 안내

원문 발행 2026. 08. 13.수집 2026. 08. 13.출처 타입 TWITTER

인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.