본문으로 건너뛰기

GPT-6 Astra 확산, Ling 의료·비전 모델, 코딩 에이전트 검증 강화

GPT-6 Astra의 장시간 에이전트 작업 확산과 Ling 모델의 전문 영역 확장, 코딩·문서·수학 검증 도구의 보강

이 요약은 AI가 원문을 분석해 생성했습니다. 정확한 내용은 원문 기준으로 확인하세요.

TL;DR

이번 기간에는 GPT-6 Astra의 유료 요금제·API·GitHub Copilot·평가 플랫폼 확장이 가장 큰 흐름을 이뤘다. Astra는 Computer Use, 비동기 도구 호출, 응답 중간 제어를 결합해 장시간 코딩과 에이전트 작업을 겨냥했고, 동시에 Ling-3.0-flash-VL과 의료 특화 Ling-3.0-flash-Sante가 시각 이해와 보건 작업으로 영역을 넓혔다. Claude Code는 대형 출력 처리, 스킬 사용량 점검, 중단 제어와 MCP·Remote Control 관련 기능을 보강했다. 수학 증명 형식화, 예측 모델 Fine-tuning, 문서 추출 평가처럼 에이전트의 실행 신뢰성과 검증 가능성을 측정하려는 흐름도 이어졌다.

𝕏 실시간 트렌드 토픽

🔥 GPT-6 Astra의 API·Codex·Copilot 확산과 장시간 에이전트 실행포스트 15

GPT-6 Astra가 Pro·Enterprise·Business Premium 사용자와 API에 먼저 배포되고 ChatGPT Work, Codex, GitHub Copilot, Agent Arena, OpenResearch로 접근 범위를 넓혔다. Computer Use, 비동기 도구 호출, 응답 중간 제어를 결합해 장시간 코딩과 복합 작업을 겨냥한 배포 흐름이다.

세부 내용 보기
  • GPT-6 Astra는 Pro·Enterprise·Business Premium 사용자와 API에서 먼저 제공되고 Plus와 Business 사용자로 순차 확장되는 방식으로 배포됐다. OpenAI Developers는 이를 Computer Use와 software engineering, visual understanding을 위한 모델로 소개했고, 사용자는 ChatGPT Work와 Codex에서 접근한다.
  • Astra는 컴퓨터 조작, 비동기 도구 호출, 응답 중간 제어를 한 흐름에 결합한다. 도구 결과를 기다리는 동안 이어지는 작업을 처리하고 실행 중 추가 지시를 반영하는 구조라서 복합적인 장기 워크플로와 재작업이 많은 engineering 작업을 겨냥한다.
  • GitHub Copilot은 Astra를 앱·CLI·코드 환경에 제공하며, 내부 테스트에서 모델이 작업 중 계획을 세우고 검증을 반복한 뒤 결과를 독립적으로 확인했다고 밝혔다. 이 과정이 이전 OpenAI 모델보다 광범위한 코딩 작업을 더 적은 단계로 처리하는 결과로 이어졌다는 설명이다.
  • Astra는 Agent Arena와 OpenResearch에도 추가됐고, Arena는 웹 검색과 파일시스템을 사용할 수 있는 장시간 에이전트 작업으로 투표 기반 평가를 진행한다. OpenResearch 게시물은 Codex CLI 0.153.1 이상에서 Astra를 선택할 수 있다고 전했다.
원문 트윗 2개 보기

📈 Ling-3.0-flash-VL과 Sante의 시각·의료 작업 확장포스트 3

Ant Ling이 시각 이해와 시각 에이전트 기능을 더한 Ling-3.0-flash-VL, 의료·보건 작업에 맞춘 Ling-3.0-flash-Sante를 공개했다. 두 모델은 범용 시각 작업과 전문 의료 추론을 각각 겨냥하며 Computrix를 통한 체험 경로도 제공한다.

세부 내용 보기
  • Ling-3.0-flash-VL은 Ling-3.0-flash를 기반으로 시각 이해와 시각 에이전트 기능을 결합했다. 시각 지각, STEM reasoning, 문서 지능, 멀티모달 에이전트 작업, frontend coding, 의료 보고서 해석을 하나의 모델 범위에 포함한다.
  • Ling-3.0-flash-Sante는 Ling-3.0-flash 기반 MoE 모델로 의료 추론, 전문 healthcare 작업, 심층 research, 근거 기반 retrieval을 처리하도록 구성됐다. MedXpertQA-Text, DiagnosisArena-MCQ, AFUMED-Drug, HealthBench Professional, BrowseComp 결과가 함께 언급됐다.
  • Ling-3.0-flash-VL은 Computrix에서 2주간 무료로 사용할 수 있고, API 접근에는 $20부터 시작하는 Token Plan이 필요하다. 이미지나 비디오 입력은 Base64 형식으로 전달하는 방식이다.
원문 트윗 2개 보기

Claude Code의 대형 컨텍스트·중단 제어·MCP 운영 보강포스트 4

Claude Code 2.1.261이 명령과 백그라운드 작업 출력 한도를 128K로 높이고, 사용하지 않는 skill과 컨텍스트 비용을 점검하는 기능을 추가했다. 중단 신호, Remote Control, MCP 서버 관리, 프록시·인증·세션 복구 관련 오류도 함께 손봤다.

세부 내용 보기
  • 이번 버전은 inline command와 background output 한도를 128K로 높여 큰 출력이 파일로 빠지기 전에 컨텍스트 안에 유지하도록 했다. /skill-doctor는 로드됐지만 사용하지 않는 skill과 컨텍스트 비용을 보여줘 불필요한 항목을 줄이는 입력 관리 흐름을 제공한다.
  • SDK와 cloud session은 첫 prompt 직후 들어온 Stop·interrupt를 반영하도록 바뀌었다. 이전에는 turn이 시작되기 전에 보낸 중단이 무시될 수 있었지만, 수정 후에는 실행을 멈추는 입력 제어로 처리된다.
  • VS Code에서는 MCP 서버를 IDE를 떠나지 않고 추가·삭제하는 대화상자가 제공되고, Remote Control 세션의 권한 상태·중단·세션 표시 문제를 수정했다. 프록시 뒤의 이벤트 스트림, connector 재시도, 세션 복구와 관련된 오류 처리도 함께 조정됐다.
  • LangChain은 MCP 지원을 기본 패키지에 포함하고 FastMCP v4 기반의 무상태 사양으로 전환했다고 밝혔다. 따라서 `uv pip install 'langchain[mcp]'`로 설치한 뒤 MCP 서버를 LangChain 흐름에 연결하는 경로가 마련됐다.
원문 트윗 2개 보기

Claude Code Changelog

@ClaudeCodeLog

2시간 전

Claude Code CLI 2.1.261 changelog: New features: • Added an "Organization policy" line to /status and claude doctor that says why your organization's policy could not be loaded, such as a proxy not passing the endpoint through • Added bashOutputMaxChars and taskOutputMaxChars settings to raise how much command and background-task output Claude receives inline before it is saved to a file, up to 128K characters • Added --append-subagent-system-prompt-file to read the subagent system prompt from a file, for prompts too large to pass on the command line • Added /skill-doctor to show which loaded skills go unused and what they cost in context, so you can prune them Fixes: • Fixed typed or pasted characters occasionally landing out of order or being dropped during fast input or key repeat • Fixed /add-dir <subdirectory> printing a false "couldn't be resolved" error when the working directory is on a /net automount • Fixed the Bedrock setup wizard hanging when AWS or an AWS credential helper never responds (it now times out with a clear error), and its model checks failing behind a TLS-inspecting proxy • Fixed cloud sessions discarding a plugin synced from http:// claude.ai when managed settings force-enable it in enabledPlugins, then falling back to a marketplace clone that could fail • Fixed being unable to delete the character immediately before an inline [Image #N] chip in the prompt input • Fixed resuming a session losing hook output and other context around parallel tool calls, which changed the resumed request • Fixed Remote Control showing a stale permission mode when a phone, browser, or http:// claude.ai app attaches to a terminal session or after the mode changes in the terminal • Fixed Remote Control sessions showing as still working (stuck spinner and Stop button) after stopping a turn from a connected phone or browser, or after a local slash command like /clear • Fixed SDK and cloud sessions ignoring a Stop or interrupt sent just after the first prompt, before the turn had started; the turn now stops instead of running to completion • Fixed Remote Control uploading a session pulled with /teleport into the connected session, which appeared appended to the original on phone and web • Fixed Remote Control's inbound event stream failing behind TLS-inspecting corporate proxies on native Windows • Fixed Remote Control sessions showing the default effort level on http:// claude.ai when the effort comes from settings • Fixed gcpAuthRefresh opening a browser at startup when the Google credential check was slow, even though the credential was still valid • Fixed http:// claude.ai connectors staying absent for the whole session when the startup connector fetch timed out — the CLI now retries in the background • Fixed sustained high CPU usage when a background agent could not be resumed and its wake-up was retried in a tight loop • Fixed feature flags gated to a newer version occasionally applying to an older Claude Code version running on the same machine • Fixed /usage and the VS Code usage panel dropping a model-specific weekly limit row when the usage endpoint is rate limited or when opened right after startup • Fixed claude -p --resume <file> adopting a malformed session ID recorded in the transcript; it now resumes under a fresh session ID instead • Fixed the terminal progress indicator (iTerm2, Ghostty, ConEmu) showing the session as finished while a background workflow or agent was still running • Fixed a rare layout glitch where a box could render with the wrong height after its container switched between row and column direction • Fixed Claude apps gateway client IP when a trusted proxy appends a port to X-Forwarded-For; with an access list set, an unreadable entry now gets 403 • Fixed Claude apps gateway telling Claude Desktop to export OpenTelemetry as JSON even when the terminal CLI uses protobuf, so protobuf-only collectors rejected Desktop's data • Fixed Desktop and web showing a session as busy while it only watches an artifact for updates • Fixed Claude in Chrome file_upload failing with "paths: expected array, received undefined" in local Cowork sessions run from the Claude Desktop app • Fixed SendMessage to an offline Remote Control session on another machine reading as delivered; the result now says delivery is queued until that machine reconnects • Fixed plugin install hints from CLIs run in background Bash commands: they are now detected, and the raw <claude-code-hint> tag no longer leaks into the conversation • Fixed in-process agent-team teammates re-sending their first-turn tool and skill announcements on the second turn, which changed the request prefix and missed the prompt cache Improvements: • Improved the /model picker and the VS Code model pill to show a model's name instead of its raw Bedrock, Vertex AI, or LLM gateway ID when Claude Code recognizes it • Improved startup on Google Vertex AI when GOOGLE_APPLICATION_CREDENTIALS is set: API client creation no longer re-runs Google Cloud project discovery or spawns extra gcloud processes • Improved streaming performance: already-rendered blocks are no longer re-checked by layout on each update • Improved the dangerous-rm safety prompt to also catch rm -rf on positional parameters and inside double-quoted sh -c scripts • Improved handling when the API sends no response headers: the retry now waits up to API_TIMEOUT_MS (10 minutes by default) instead of another 3 minutes, and the messages say what to change Security/safety changes: • [VSCode] Added a fold button to permission and question prompts so the conversation behind them can be read without dismissing them; the space beside the prompt now scrolls the conversation • [VSCode] Fixed the next queued permission prompt keeping text typed on the previous prompt and accepting an immediate second click Other changes: • Changed a Claude apps gateway 403 on the managed settings load (at startup or after /login) to say Claude Code may not be enabled for the organization, instead of advising a new sign-in • Changed machines whose managed settings pin forceLoginMethod: "gateway" to ignore a leftover API key or http:// claude.ai login and ask for /login; Bedrock, Vertex AI, and Foundry sessions are unaffected • Changed auto mode to treat a link that packs content into a public diagram renderer's URL as an upload to that site: no longer auto-approved unless you asked for it • Changed the prompt's word-editing keys to match Bash: Ctrl+W deletes back to whitespace, Alt+F and Alt+D stop at word end, punctuation separates words; keybindingFlavor no longer has any effect • Changed /context token counting to use a local estimate when the token-counting API is unavailable, instead of extra small-model requests • [VSCode] Added a "Build a custom style" walkthrough to the Output styles menu that writes a custom output style file and lists it right away • [VSCode] Added an Add server form and a Remove action to the MCP servers dialog, so MCP servers can be added and removed without leaving the IDE • [VSCode] Added a hollow ring in the session list for sessions open in a terminal, another VS Code window, or Claude Desktop, so they no longer look closed • [VSCode] Added "Archive session" to the session list's right-click menu and gave Unarchive its own icon • [VSCode] Fixed a session teleported from Claude Code on the web treating a question that was cut off when the cloud session shut down as declined • [VSCode] Fixed the session tab's Rename box opening empty for a tab restored with the window; it now starts with the current name • [VSCode] Fixed collapsed sections in the session list panel briefly showing expanded each time the panel loaded • [VSCode] Fixed Focus view showing a tool call as still running after Claude had moved on, such as while a question waited for your answer • [VSCode] Fixed the session list's active-row highlight going stale when an unfocused Claude tab's session ID is corrected • [VSCode] Fixed Cmd/Ctrl+Shift+T reopen and deep-link opens placing the Claude tab outside the Claude editor group when a Claude tab has focus • [VSCode] Fixed the session tab's "Add to group" putting a session opened from Claude Code on the Web in two groups; it now moves the entry the session list shows • [VSCode] Fixed the model picker showing models an organization has since disabled until the window was reloaded twice • [VSCode] Fixed a tab opened from the session list jumping back to that session, and a tab opened from a Web session restarting its teleport or staying empty, after VS Code reloads the tab's view • [VSCode] Fixed /btw side-question history from earlier sessions being overwritten when a question is asked right after a window reload or while a settings file has errors • [VSCode] Fixed the pending question card not reappearing after the Claude panel reloads when signed in with a http:// Claude.ai or Console account • [VSCode] Fixed claude.ai-only features staying visible in a window's other Claude panels after one panel picked up a third-party provider from a settings file • [VSCode] Fixed the sign-in screen appearing despite the Disable Login Prompt setting when Claude Code reports no login or a request fails for lack of one • [VSCode] Fixed install-plugin links opening the Claude sidebar without the install dialog in a window where only the session list had been shown • [VSCode] Fixed the sidebar usage meter staying empty on a new window until the Account & usage dialog was opened, and a 0% usage limit being left out of the meter • [VSCode] Fixed "Start new session in this group" losing the group after New conversation, and a missing unread dot for a session that finished before the sidebar's unread list loaded • [VSCode] Fixed the editor tab badge showing unread during a running turn or missing on a tab opened from the session list, and "Add Session Tab to Group" doing nothing for an archived session • [VSCode] Fixed "Enable Remote Control for all sessions" so flipping it also applies right away to sessions open in other VS Code windows • [VSCode] Fixed the session list's Open filter for sessions continued from http:// claude.ai whose tab was still recorded under the web session, and labeled the filter menu's sections for screen readers • [VSCode] Changed the model picker to one flat list of every model, with rows kept for older model spellings listed last Source: https:// github.com/anthropics/cla ude-code/blob/main/CHANGELOG.md#21261 …

💬 1 0 0👁 140

LangChain OSS

@LangChain_OSS

4시간 전

MCP usage is growing... and growing fast! check out our latest blog on our revamped support for MCP and the new, stateless spec!

Sydney Runkle

ICYMI -- yesterday we released support for the new MCP protocol in @LangChain !! what do you need to know? 1. MCP support is now in the main langchain package! get started with `uv pip install 'langchain[mcp]'` 2. it's now built on top of FastMCP v4; FastMCP has ergonomic x.com/sydneyrunkle/s…

💬 0 1 3👁 159

📈 수학 증명 형식화와 에이전트 추론 신뢰성포스트 3

Anthropic은 Fermat’s Last Theorem의 형식화된 증명 사례를 통해 Lean 같은 증명 보조기의 기계 검증 흐름을 전했다. 별도 게시물에서는 실행 가능한 코드만 학습한 로컬 Coding 모델과 에이전트 작업용 벤치마크가 추론 결과의 실제 성공 여부를 가르는 기준으로 떠올랐다.

세부 내용 보기
  • 수학적 증명의 정확성을 사람이 확인하는 데는 수년이 걸릴 수 있지만, Formalization은 추론을 Lean 같은 proof assistant가 검사할 수 있는 형태로 바꾼다. Anthropic은 Claude가 지난달 Fermat’s Last Theorem의 첫 형식화된 증명을 완성했다고 전했다.
  • Gemma 4 12B Coder는 실행 후 테스트를 통과한 Python CoT만 학습 자료로 남기는 방식으로 구축됐다고 소개됐다. Q2_K 형식은 4.83GB이며 Q4는 12GB 카드와 Apple Silicon에 맞고, 최대 256K 컨텍스트를 제공한다는 수치가 함께 제시됐다.
  • Agent Arena는 웹 검색과 파일시스템을 사용할 수 있는 실제 장기 에이전트 작업을 대상으로 모델을 측정하고, Terminal-Bench-Science 0.1 관련 게시물은 GPT-6 Astra가 Fable 5.1을 앞선다고 전했다. 계산 과정의 그럴듯함보다 실행 결과와 독립 검증을 평가 기준으로 삼는 흐름이다.
원문 트윗 2개 보기

📈 예측 Fine-tuning과 문서 추출 평가의 실전 데이터화포스트 2

Tinker는 참조 정보를 바탕으로 사건 확률을 예측하는 모델을 Fine-tuning하는 cookbook과 ProphetArena 데이터셋을 공개했다. Kaggle과 LlamaIndex는 누락값·근거·반복 구조 보존을 측정하는 ExtractBench를 내놓아 문서 기반 에이전트의 오류 유형을 구체화했다.

세부 내용 보기
  • Tinker의 cookbook은 참조 정보를 입력으로 받아 사건 확률을 예측하도록 모델을 학습하는 Fine-tuning 절차를 제공한다. 사용자는 ProphetArena가 제공한 데이터셋으로 시작한 뒤 자체 데이터로 확장하며 예측 모델의 학습 흐름을 재현한다.
  • ExtractBench는 공급망, healthcare, finance 문서를 대상으로 schema-guided extraction 뒤 사람 검토가 이뤄지는 작업을 측정한다. 시스템은 없는 값을 임의로 만들지 않고 null로 반환해야 하며, 각 값의 출처 근거와 반복 구조의 모든 레코드를 보존해야 한다.
  • 공개 게시물은 GPT-5.6 Sol이 ExtractBench에서 91%로 선두라고 전했다. 단순한 답변 정확도보다 누락값 처리, 출처 연결, 구조 보존을 함께 측정해 실제 결제와 의사결정에 연결되는 오류를 평가하는 방식이다.
원문 트윗 2개 보기

에이전트 실행을 받치는 Sandbox·Trajectory·패션 제작 파이프라인포스트 3

Perplexity는 Computer를 뒷받침하는 Rust 기반 Sandbox 플랫폼을 공개 행사에서 다룰 예정이며, LangSmith 사례는 디자인 요청을 렌더·tech pack·캠페인 이미지로 바꾸는 LangGraph 흐름을 전했다. 에이전트의 실행 격리와 작업 기록 관리가 제품 기능의 기반으로 묶이는 모습이다.

세부 내용 보기
  • Perplexity는 Perplexity Computer의 Sandbox를 Rust로 구축한 SPACE 플랫폼을 RustConf에서 다룬다. Sandbox는 에이전트가 수행하는 작업을 별도 실행 환경에서 처리하는 기반으로 언급됐고, 발표 주제는 그 플랫폼의 내부 구조다.
  • Raspberry AI는 디자인 팀의 보드와 함께 작동하며 자연어 요청을 완성된 render, tech pack, campaign imagery로 변환하는 패션용 agentic platform이다. 전체 LangGraph agent lifecycle을 LangSmith에서 실행하는 구조라서 요청 입력부터 결과 생성과 운영 기록까지 하나의 흐름으로 묶인다.
  • LangChain은 agent trajectory를 처음부터 저장하기 위한 smithdb를 구축했다고 밝혔다. 에이전트가 어떤 도구 호출과 상태 변화를 거쳐 결과에 도달했는지 기록하는 저장 계층이 장기 실행 시스템의 운영 기반으로 부상한 사례다.
원문 트윗 2개 보기

용어 해설

컴퓨터 사용(Computer Use)
모델이 텍스트 응답에 그치지 않고 화면과 애플리케이션을 읽으며 클릭, 입력, 파일 조작 같은 컴퓨터 작업을 수행하는 방식이다. GPT-6 Astra는 비동기 도구 호출과 응답 중간 제어를 결합해 장시간 작업을 처리하는 데 활용한다.
비동기 도구 호출(Asynchronous Tool Calling)
모델이 도구 실행 결과를 기다리는 동안 다른 작업 흐름을 이어가고, 결과가 도착하면 처리에 반영하는 호출 방식이다. 긴 작업에서 대기 시간을 줄이고 여러 단계를 연결하는 데 쓰인다.
응답 중간 제어(Mid-response Steering)
모델이 응답을 완전히 끝내기 전에 사용자의 추가 지시나 실행 결과를 받아 현재 작업의 방향을 조정하는 방식이다. GPT-6 Astra는 이를 Computer Use와 도구 호출에 결합한다.
형식화(Formalization)
수학적 추론을 Lean 같은 컴퓨터 증명 보조기가 검사할 수 있는 형식으로 변환하는 과정이다. 자연어 증명의 각 단계를 기계가 검증할 수 있게 만들어 대규모 정리의 정확성 확인에 쓰인다.
전문가 혼합 모델(MoE)
하나의 모델 안에 여러 전문가 네트워크를 두고 입력마다 일부 전문가를 선택해 계산하는 구조다. Ling-3.0-flash-Sante는 의료·보건 작업을 겨냥한 MoE 모델로 소개됐다.
모델 컨텍스트 프로토콜(MCP)
AI 애플리케이션과 외부 도구·서버를 연결하는 프로토콜이다. LangChain은 MCP의 무상태 사양을 지원하고, MCP 서버를 패키지와 IDE에서 연결하는 흐름을 확장했다.
AI 분석 전체 내용 보기

AI 요약 · 북마크 · 개인 피드 설정 — 무료

출처 · 인용 안내

원문 발행 2026. 09. 05.수집 2026. 09. 05.출처 타입 TWITTER

인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.