TL;DR
이번 기간에는 X 추천 알고리즘의 소스 공개, Grok 4.6의 벤치마크와 비용 비교, ChatGPT의 Computer History가 큰 비중을 차지했다. Grok 4.6은 GPQA Diamond에서 94.9%를 기록했다는 포스트와 WANDR에서 Fable 5와 같은 결과를 60% 이상 낮은 비용으로 냈다는 포스트가 함께 확산됐다. Gemini 3.7 Flash와 GPT-5.6 Sol은 속도와 가격을 앞세웠고, Claude 기반 앱 유지보수와 Hermes Agent·Perplexity Agent API 같은 에이전트 실행 구조도 이어졌다. MiniMax Music 3의 로컬 공개와 에이전트 기반 ICML 논문 재현은 모델 공개 범위와 자동화된 검증 방식의 확장으로 이어졌다.
𝕏 실시간 트렌드 토픽
🔥 X 추천 알고리즘 소스 공개와 투명성 검증포스트 4
X가 For You 추천 알고리즘의 소스 코드를 공개하고 외부 비판과 개선 의견을 받는 흐름이 확산됐다. 공개된 변경 내역과 계정 노출 제한 보고서가 추천 결과를 확인하는 경로로 연결됐다.
- X의 추천 결과가 왜 달라졌는지 외부 이용자가 확인하기 어려웠던 상황에서 For You 알고리즘 코드와 변경 문서가 공개됐다. 공개 코드는 게시물에 별도의 품질 점수를 부여하기보다 각 이용자가 게시물에 취할 행동을 추정하고, 행동별 예측값에 가중치를 곱해 추천 순서를 계산하는 구조로 요약됐다. kcoleman은 친구 게시물이 타임라인에 많이 나타났던 변경도 BIDIRECTIONAL_BOOST_CHANGE 문서에서 확인할 수 있다고 전했다. 소스와 변경 기록이 함께 공개되면서 추천 결과를 감사하고 비판할 수 있는 범위가 넓어진 셈이다.
- cb_doge는 계정이나 게시물이 노출 제한을 받았는지, 그 이유가 무엇인지 확인하고 전체 보고서를 내려받을 수 있다고 전했다. 이는 알고리즘 코드 공개가 내부 구현 공개에만 머물지 않고 개별 계정의 가시성 판단을 확인하는 사용자 경로와 결합된 사례다.
원문 트윗 2개 보기

Elon Musk
@elonmusk
We are making 𝕏 open source. Transparency build trust.
“Am I shadowbanned?” “Is X fair?” “Why am I seeing this post?” We want the public to be able to answer these themselves. Today we’re giving an unprecedented level of transparency into the X algorithm so people can audit, critique and even help improve it. Enjoy & send feedback. x.com/XOpenSource/st…
@kcoleman
You know how everyone started seeing their friends in the timeline a few weeks ago? With our new open-source code base, you'd also have been able to see what changed under the hood -- here's the story: https:// github.com/xai-org/x-algo rithm/blob/main/docs/BIDIRECTIONAL_BOOST_CHANGE.md …
🔥 Grok 4.6의 벤치마크 점수와 작업당 비용 경쟁포스트 6
Grok 4.6이 GPQA Diamond, RareBench, WANDR 관련 포스트에서 높은 점수와 비용 효율을 함께 내세웠다. Perplexity와의 연동으로 연구형 작업에 투입되는 경로도 확장됐다.
- Grok 4.6은 GPQA Diamond에서 94.9%로 1위에 올랐다는 포스트가 확산됐고, GPT-5.6·Gemini 3.1 Pro·Claude Opus 5와 비교됐다. RareBench 관련 포스트에서는 어린이 희귀질환 진단에서 Claude Opus 5를 앞섰으며 비용은 약 3분의 1이라는 수치가 제시됐다. 벤치마크별 평가 대상과 비용 기준은 서로 다르지만, 단순 점수보다 전문 작업 성능과 지출을 함께 묶는 비교 방식이 반복됐다.
- Perplexity의 WANDR에서는 Grok 4.6이 Claude Fable 5와 같은 0.496점을 기록하면서 작업당 비용은 7.58달러 대 20.30달러로 제시됐다. Perplexity와 Perplexity Computer에 모델이 연결되면서 다단계 연구 작업의 결과와 실행 비용을 한 화면에서 비교하는 사용 경로가 생겼다.
Grok 4.6은 여러 벤치마크에서 경쟁 모델과 비슷하거나 높은 결과를 내면서 작업당 비용을 크게 낮췄다는 평가를 받았다.
원문 트윗 2개 보기

Elon Musk
@elonmusk
Try Grok 4.6
Grok 4.6 wins again. Grok 4.6 takes the #1 spot on GPQA Diamond with a score of 94.9%, beating GPT-5.6, Gemini 3.1 Pro, Claude Opus 5, and every other model tested by Artificial Analysis.

Aravind Srinivas
@AravSrinivas
Congrats to @SpaceXAI on one more amazing model: Grok 4.6. We benchmarked it as an orchestrator on our Wide-And-Deep-Research benchmark using the Perplexity Computer harness, and it neatly sits on the Pareto frontier of performance vs cost. Available to all Pro and Max users on Perplexity!
Grok 4.6 is now available in Perplexity and Perplexity Computer. On WANDR, it sits on the Pareto frontier of performance and efficiency, matching Fable 5 results at over 60% lower cost.
📈 Gemini 3.7 Flash와 초고속 모델의 가격·지연 경쟁포스트 5
Gemini 3.7 Flash는 이전 세대보다 낮은 가격과 짧은 출시 주기의 성능 개선을 내세웠다. GPT-5.6 Sol의 Ultrafast 모드까지 겹치며 에이전트용 하위 모델의 속도와 비용이 함께 비교됐다.
- Gemini 3.7 Flash는 3.6 Flash 출시 약 3주 뒤 등장했고, 연말까지 가격을 50% 낮추며 코딩과 agentic work에서 지능 향상을 제공한다는 포스트가 나왔다. API, AI Studio, Antigravity 등 여러 경로에 배포된다는 내용도 함께 제시됐다. Perplexity는 Flash 모델을 다중 모델 하네스 안의 빠르고 비용 효율적인 subagent로 사용한다고 전했다.
- OpenAI가 GPT-5.6 Sol의 Ultrafast 모드를 최대 14배 속도로 제공한다고 밝히면서 속도 경쟁의 기준이 더 짧은 응답 시간으로 이동했다. 한쪽은 50% 가격 인하와 약 3주 만의 세대 교체를, 다른 쪽은 최대 14배 속도를 앞세워 반복 호출이 많은 에이전트 작업의 비용과 지연을 동시에 겨냥했다.
원문 트윗 2개 보기

Josh Woodward
@joshwoodward
3.7 Flash: Fast, 50% cheaper, happened in ~3 weeks
Introducing Gemini 3.7 Flash : ) - it is fast! - 50% lower price than 3.6 flash (through end of year) - strong intelligence increase in only ~3 weeks (thanks to some awesome algorithmic improvements) - available in the API, AI Studio, Antigravity, and more!
Tibo
@thsottiaux
Sometimes you have have to go /ultrafast.
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows.
📈 ChatGPT Computer History의 작업 맥락 기억포스트 6
ChatGPT 데스크톱 앱에 최근 앱·웹사이트 활동을 기억하는 Computer History 기능이 배포됐다. 사용자는 기록 범위와 일시정지 여부를 조절하고, 반복 작업에서 더 풍부한 맥락을 제공할 수 있다.
- Computer History는 사용자가 컴퓨터에서 수행한 최근 작업을 타임라인으로 남기고 ChatGPT가 다음 상호작용에서 그 맥락을 활용하는 흐름이다. 사용자는 Settings → Integrations에서 기능을 켜고, 앱과 웹사이트를 포함하거나 제외하며, 전체 또는 일부 기록을 지우고, 기능을 일시정지하거나 다시 시작할 수 있다. OpenAI는 Chronicle research preview를 바탕으로 토큰 사용량을 줄이고 privacy controls를 추가했다고 밝혔다.
- Codex와 ChatGPT는 최근 작업 맥락을 바탕으로 중단된 지점에서 이어가고 반복되는 업무 패턴을 파악해 skill이나 예약 작업을 제안할 수 있다고 전했다. Pro·Business·Enterprise 사용자를 대상으로 Mac 데스크톱 앱에서 전 세계 배포가 진행되며, EEA·영국·스위스는 이후 수 주 안에 접근이 예정됐다.
작업 기억은 반복 설명을 줄이고 연속 작업을 이어가는 데 유용하지만, 앱·웹사이트 기록의 포함 범위와 삭제·일시정지 설정을 사용자가 직접 관리해야 한다.
원문 트윗 2개 보기
OpenAI
@OpenAI
ChatGPT can now remember your activity across the apps and websites on your computer. With Computer History in the desktop app, future interactions feel more personalized and require less explanation.
OpenAI
@OpenAI
Computer History builds on the Chronicle research preview with reduced token usage and more privacy controls. A new timeline view gives you the ability to look back on your work and build skills from your frequent tasks. From there or the menu bar, you can: - Clear all or parts of your history - Include or exclude apps and websites - Pause and resume Computer History
📈 Claude 기반 앱 유지보수와 코딩 에이전트의 실행 루틴포스트 3
Claude Code와 Claude Tag를 이용해 충돌 탐지, 중복 추상화 통합, dead code 제거를 매일 실행하는 앱 유지보수 흐름이 공유됐다. Cursor의 Firetiger 인수와 Claude Code의 자동 재개 기능도 코딩 에이전트의 운영 범위를 넓혔다.
- 앱 유지보수용 Slack 채널에서 Claude Tag가 iOS·Android·Desktop·web·CLI·Agent SDK를 대상으로 crash fuzzer, dup unifier, dead-code remover, abstraction police 같은 루틴을 매일 실행한다. 시뮬레이터에서 충돌을 찾고 원인을 수정하거나, 코드베이스의 유사 추상화를 찾아 PR을 만들고, 정적으로 도달할 수 없는 코드를 제거하는 입력·처리·출력 흐름이다. 몇 주 동안 388개 PR이 열렸고 그중 180개가 Claude Code Review와 human review 뒤 병합됐다는 수치가 제시됐다.
- 이 방식은 에이전트가 한 번 코드를 생성하는 데서 끝나지 않고 실행 결과를 바탕으로 다음 날 루틴을 조정하는 운영 구조다. Claude가 첫 시도에 맞히지 못하면 루틴을 튜닝해 반복 개선하고, Claude Code 데스크톱의 auto-continue checkbox는 사용량 제한이 풀린 뒤 중단된 작업을 자동으로 이어간다. Cursor는 Firetiger 팀과 함께 작업을 production까지 따라가 오류를 고치는 agent 구축을 추진한다고 밝혔다.
원문 트윗 2개 보기
Boris Cherny
@bcherny
A weird experiment I've been trying the last few weeks is having Claude take over day-to-day maintenance of our apps. Seeing early signs of life that this might be possible. The setup is straightforward: we have a Slack channel called proj-claude-maintains-apps. In it, Claude Tag runs a bunch of daily routines across iOS, Android, Desktop, web, CLI, and Agent SDK: - Crash fuzzer: open the app in a simulator and tap around to find ways to crash it, then root cause and fix the crashes - Dup unifier: scans the codebase for similar-yet-slightly-divergent abstractions, and puts up PRs to unify them - Dead-code remover: removes statically unreachable code, and adds logging to suspected dead code to check if it's really dead and if so, remove it the next day - Abstraction police: fixes leaky abstractions - a bunch more.. Results have been surprisingly positive. Over the last few weeks, these routines have opened 388 PRs across our repos, 180 of which we merged after Claude Code Review + human review. We're now thinking about how to streamline this to make merging these kinds of mechanical changes easier. Claude generally gets these PRs right on the first shot, and if it doesn't, we ask Claude to tune its routines so it's better the next day. Sometimes it takes a few days of tuning. To try a similar workflow, ask Claude Code or Tag, or create some routines directly at https:// claude.ai/code/routines. A few of the actual prompts I used below. Has anyone experimented with similar workflows?

ClaudeDevs
@ClaudeDevs
Hit your usage limit in Claude Code desktop? There's now an auto-continue checkbox. Turn it on, and it'll automatically continue where you left off once your limit resets.
📈 에이전트의 다중 실행과 연구용 API 확장포스트 5
Hermes Agent는 여러 bot이 서로 통신하고 하위 에이전트의 실시간 transcript를 제어하는 기능을 추가했다. Perplexity Agent API와 Managed Deep Agents는 검색·코드 실행·예약 실행을 하나의 작업 흐름으로 묶는 방향을 보였다.
- Hermes Agent의 Bot Mode는 세션마다 하나의 대화를 두는 대신 agent profile별 bot을 만들고, 각 bot에 job·description·profile picture를 부여한 뒤 서로 통신하게 한다. 별도 기능은 하위 에이전트가 비동기로 작업하는 동안 live transcript를 읽고, 멈추고, 조종하는 제어 경로를 제공한다. 공개 beta를 거쳐 Hermes Desktop의 기본 앱에 반영할 계획이라는 내용이 나왔다.
- Perplexity Agent API는 grounded web search를 유지하면서 multi-step research, code execution, built-in tools, multiple models를 한 API에 결합하고 BrowseComp와 WideSearch에서 기존 Sonar 최고 점수를 두 배 이상 웃돈다고 전했다. Managed Deep Agents는 schedule을 부여해 스스로 prompt를 생성하는 에이전트를 몇 줄의 코드로 구성하고, 이 흐름이 단일 채팅을 넘어 예약 실행과 도구 조율로 확장되는 구조다.
원문 트윗 2개 보기

Teknium
@Teknium
Introducing Bot Mode for Hermes Agent. Bot Mode is an alternative to sessions mode, where you have one chat with each agent profile, or "bot". These bots can be given jobs, descriptions, profile pics, and communicate with your other bots! For one day we will do a public beta test of Bot Mode for Hermes Desktop through this plugin: https:// github.com/NousResearch/H ermes-Bot-Mode … Give me all your feedback here on how it does for you, if it works as expected, any bugs. Then we'll address the feedback and get it into the main Desktop App for everyone!
봇 모드: Hermes 에이전트 데스크톱 앱에 출시 예정인가요?

Perplexity Developers
@perplexitydevs
Sonar is moving to the Agent API. The Perplexity Agent API keeps grounded web search, and adds multi-step research, code execution, built-in tools, and access to multiple models through one API. On BrowseComp and WideSearch, Agent API more than doubles the best Sonar score.
📈 오픈 음악 모델의 로컬 실행과 에이전트 기반 논문 재현포스트 3
MiniMax Music 3가 ComfyUI에서 로컬로 실행되도록 공개됐고, ICML 논문 재현에는 coding agent가 투입됐다. 모델 가중치 공개와 연구 검증 자동화가 각각 창작·학술 작업의 실행 경로를 바꾸는 사례로 묶였다.
- MiniMax Music 3는 ComfyUI에서 로컬 실행이 가능해졌고 cloud 제공은 이후로 예고됐다. 포스트에는 MiniMax가 AGI가 도달할 때까지 open 상태를 유지하겠다는 문구가 포함됐다. 기존 음악 생성 서비스의 이용 제한과 대비되면서, 사용자가 모델을 원격 서비스가 아닌 로컬 환경에서 실행하는 경로가 부각됐다.
- Hugging Face와 askalexiv의 hackathon에서는 1,200명이 ICML 2026 논문 중 3분의 1을 coding agent로 재현했다는 포스트가 나왔다. 검토된 논문 가운데 약 절반에서 검증된 내용이 발견됐고, 4분의 1에서는 반증되거나 이견이 있는 내용이 나왔으며, 일부 반증은 중심 bound가 잘못된 사례까지 포함했다. 재현 결과 자체도 틀릴 수 있어 최종 판별에는 사람이 필요했지만, 학회 심사의 첫 단계에 agent 기반 재현을 둘 수 있다는 실무적 가능성이 제기됐다.
에이전트 기반 재현을 심사 초기에 적용하면 논문 주장과 코드 실행 결과를 더 일관되게 대조할 수 있다는 입장이다.
자동 재현은 반증 후보를 찾는 데 쓰일 수 있지만 재현 과정 자체의 오류를 가려내려면 사람의 판정이 계속 필요하다는 관점이다.
원문 트윗 2개 보기
@MiniMax_AI
Music 3 is open too!! Already live locally on @ComfyUI "We will keep open until AGI arrives." Cloud coming soon.
Joël Niklaus
@joelniklaus
Stop what you're doing and read this blog post now, this is such a cool project! TLDR: @HuggingFace and @askalphaxiv ran a hackathon where 1,200 people pointed coding agents at a third of ICML 2026 and tried to reproduce every paper claim by claim. About half of the examined papers had something verified, a quarter had something falsified or contested, and a handful of the falsifications are serious, including a spotlight paper whose central bound is simply wrong. The reproductions themselves were sometimes wrong too, and sorting that out still took humans. I think we should have automated agent based reproductions as a default first step in the conference reviewing process!
We wrote up a blog post about what we learned by reproducing 2,200 accepted ICML papers with agents, including the falsifications that we found https:// huggingface.co/blog/icml-2026 -open-reproductions …
용어 해설
- 오픈소스 알고리즘(Open-Source Algorithm)
- — 추천 알고리즘의 소스 코드를 공개해 입력값과 처리 규칙을 외부에서 확인하고 수정 제안을 낼 수 있게 하는 방식이다. 이번 포스트에서는 X의 For You 알고리즘 공개와 투명성 확보가 핵심으로 다뤄진다.
- 파레토 프런티어(Pareto Frontier)
- — 두 가지 평가축에서 한쪽을 더 개선하려면 다른 쪽을 희생해야 하는 지점들의 경계다. Grok 4.6 관련 포스트에서는 성능과 비용을 함께 비교해 효율이 높은 모델 위치를 나타내는 기준으로 쓰였다.
- Computer History
- — ChatGPT 데스크톱 앱이 사용자가 컴퓨터에서 최근 수행한 앱·웹사이트 활동을 기억하도록 하는 기능이다. 사용자가 기록 범위를 조절하거나 전체 또는 일부 기록을 삭제할 수 있고, 반복 작업을 위한 기술 구축에도 활용된다.
- Agent API
- — 웹 검색, 다단계 조사, 코드 실행, 내장 도구, 여러 모델 호출을 하나의 API 흐름으로 묶는 인터페이스다. Perplexity의 Agent API는 Sonar의 기존 검색 기능을 유지하면서 복합 작업 처리 경로를 확장한다.
- 성능·비용 효율(Performance-Cost Efficiency)
- — 동일하거나 유사한 작업 결과를 얻는 데 필요한 비용과 성능을 함께 비교하는 관점이다. WANDR와 GPQA Diamond 관련 포스트에서는 모델의 점수만 따로 보지 않고 작업당 비용 또는 속도와 결합해 평가한다.
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.