TL;DR
이번 기간에는 기업용 AI의 성능 확대보다 데이터 통제와 실행 비용을 함께 관리하는 제품 구조가 두드러졌습니다. OpenAI는 Zero Data Retention을 유지하면서 관련 상호작용의 위험 패턴을 자동 점검하는 Private Safety Processing을 시험하고 있으며, Databricks는 500페이지·100만 토큰을 넘는 문서를 병렬 분해와 결과 조정으로 처리하는 AI Extract를 내놓았습니다. Agent Harness와 Model Gateway는 도구 호출·컨텍스트·모델 라우팅을 제어해 에이전트 비용과 제공자 종속을 줄이는 계층으로 부상했습니다. 동시에 학생용 AI 요금제, 로컬 오픈 모델, 음성·영상 생성 어댑터가 사용 범위를 넓히고 있습니다.
𝕏 실시간 트렌드 토픽
📈 Zero Data Retention을 유지하는 안전 처리포스트 3
OpenAI가 장기적이고 자율적인 AI 작업의 위험 패턴을 관련 상호작용 단위로 찾으면서도 고객 콘텐츠와 데이터 통제를 유지하는 Private Safety Processing을 예고했습니다.
- 기존 Zero Data Retention 배포에서는 안전 시스템이 여러 상호작용에 걸친 위험 신호를 확인하기 어려웠고, Private Safety Processing은 고객 인프라에 콘텐츠를 둔 채 자동화 시스템이 관련 요청의 패턴을 찾아 제한된 안전 신호만 반환하는 방식입니다. OpenAI 직원에게 underlying prompts나 responses를 노출하지 않으며, 고객 제어 키로 암호화한 OpenAI 호스팅 방식도 개발 중이라는 내용이 함께 제시됐습니다. 이를 통해 장기적이고 자율적인 작업의 안전 점검과 민감 데이터 통제를 같은 배포 조건 안에서 맞추려는 구조입니다.
원문 트윗 2개 보기
OpenAI
@OpenAI
We will continue to offer Zero Data Retention for frontier models. As AI takes on longer, more autonomous work and delivers greater value to businesses, safety systems also need to identify risks across related interactions. To help address those risks, we're previewing Private Safety Processing, which is designed to improve safety without giving OpenAI personnel access to the underlying content.
Tibo
@thsottiaux
Today we’re previewing Private Safety Processing, designed to let us keep offering Zero Data Retention while improving our safeguards. Even when benefiting from frontier intelligence, customers shouldn’t have to give up control of sensitive data. For ZDR deployments, content stays on infrastructure the customer controls. Automated systems look for patterns across related interactions and return limited safety signals, without exposing the underlying prompts or responses to OpenAI employees (even me!). We’re also developing an OpenAI-hosted option encrypted with customer-controlled keys. We’re testing this with early customers now and plan to begin rolling it out in September.
We will continue to offer Zero Data Retention for frontier models. As AI takes on longer, more autonomous work and delivers greater value to businesses, safety systems also need to identify risks across related interactions. To help address those risks, we're previewing
📈 에이전트 하네스와 장기 작업 제어포스트 7
Cursor의 장기 목표·도구 호출 제어, NVIDIA SkillEvaluator를 활용한 Hermes 검증, TrueForge의 모델 중립형 실행 계층이 에이전트 신뢰성과 비용의 핵심 계층으로 묶였습니다.
- 에이전트가 긴 세션에서 목표를 유지하고 도구 호출을 끝까지 이어가려면 모델 자체뿐 아니라 실행 루프와 컨텍스트 관리가 필요합니다. Cursor는 이벤트에서 작업을 받고 목표가 완료될 때까지 진행하도록 했고, TrueForge는 도구 호출 루프·컨텍스트·하위 에이전트·샌드박스 코드 실행을 관리하며 OpenAI·Anthropic·Google과 오픈 가중치 모델을 선택할 수 있게 했습니다. 14개 기업용 에이전트 벤치마크에서 같은 Opus 4.8 답변 기준 약 30% 낮은 실행 비용, GLM-5.2 라우팅 기준 약 75% 낮은 비용이 제시됐습니다.
- Hermes는 SkillEvaluator로 설치 전 PII, 유출된 비밀값, Unicode smuggling, 라이선스와 보안 문제를 검사하고 자체 번들 스킬 11개를 개선했습니다. NVIDIA가 검증한 스킬 300개 이상을 같은 모델·같은 설정에서 비교한 결과, 스킬 사용 시 정확성이 41포인트, 효과성이 39포인트 높아졌다는 인용이 붙었습니다. 스킬을 추가하는 단계 자체를 검증 대상으로 삼아 에이전트 실행 품질과 공급망 위험을 함께 관리하는 흐름입니다.
- Cursor는 /goal로 장기 목표를 설정하고 다음 도구 호출을 기다리는 Steering 동작을 적용해 에이전트를 중간에 끊지 않도록 했습니다. Claude Code 2.1.236은 새 세션의 기본 모델 지정, macOS 샌드박스 경로 제한, 세션 간 유휴 알림을 추가했고, Google AI Studio는 GitHub 저장소 가져오기와 양방향 push/pull 동기화를 지원합니다. 목표 유지·권한 제어·저장소 작업이 개별 기능이 아니라 지속 실행 환경의 구성 요소로 결합되는 모습입니다.
원문 트윗 2개 보기

Nous Research
@NousResearch
Hermes now leverages NVIDIA's SkillEvaluator on skill installs, checking for PII, leaked secrets, Unicode smuggling, licensing and security issues before you confirm. We pointed it at our own bundled skills first, and used what it found to improve 11 of them.
We benchmarked 300+ NVIDIA verified skills to see how much they actually help agents on real tasks. Same task, same model, same setup. The only difference was whether the agent had the skill. Across the benchmarks, skills improved correctness by 41 points, effectiveness by 39,

Cursor
@cursor_ai
Use /goal to give the agent a long-lived objective to work towards until it's fully complete. Read the full changelog: https:// cursor.com/changelog/08-1 9-26 …
📈 100만 토큰 장문서 추출을 위한 전용 구조포스트 2
Databricks의 AI Extract가 500페이지를 넘고 100만 토큰을 초과하는 문서를 대규모 추출 작업으로 처리하는 맞춤 모델과 에이전트 하네스를 제시했습니다.
- 기존 문서 처리의 병목은 500페이지 이상, 100만 토큰 초과 문서와 1,000개가 넘는 객체를 포함한 중첩 스키마, 복잡한 추론이 필요한 스키마였습니다. AI Extract는 실제 기업 문서로 학습한 사내 맞춤 모델을 대량 서빙에 맞추고, 장문서 추출용 에이전트 하네스가 전체 작업을 작은 단위로 분해해 병렬 실행한 뒤 결과를 하나의 구조화된 출력으로 조정합니다. 고객의 큰 업무를 기준으로 모델과 실행 계층을 함께 만들고 연구·제품 협업으로 제공하는 방식입니다.
원문 트윗 2개 보기
Ivan Zhou
@ivanzhouyq
We are excited to share @databricks 's new AI Extract achieves the new frontier at complex document processing tasks! It addresses a few key customers' pain points: - Process documents with 500+ pages that exceeds 1M tokens - Large, nested schema with 1k+ of objects - Complex schema that requires frontier reasoning Here is how we achieved it: - In-house custom model for document extraction that is trained for complex real-world examples and optimized for serving large volume of workloads - Custom agent harness that is designed for long document extraction. It decomposes large extraction jobs, executes smaller tasks in parallel, and reconciles them into one final structured output. This is the methodology behind many of our work: identify key large workloads of enterprise customers, build in-house models + harness to address it, and ship to the customers through a close collaboration between the research + product!

Matei Zaharia
@matei_zaharia
Some continued great results with our custom trained models at Databricks!
We are excited to share @databricks 's new AI Extract achieves the new frontier at complex document processing tasks! It addresses a few key customers' pain points: - Process documents with 500+ pages that exceeds 1M tokens - Large, nested schema with 1k+ of objects - Complex
📈 여러 모델 제공자를 묶는 비용·지연시간 라우팅포스트 2
Router와 Model Gateway 활용 사례가 요청마다 작업에 맞는 모델을 선택하고 제공자별 비용·토큰 사용량을 통제하는 구조를 부각했습니다.
- 모델 제공자마다 비용, 지연시간, 추론 성능, 서비스 등급과 오픈·클로즈드 모델의 선택지가 계속 달라지는 상황에서 단일 제공자 고정은 교체 비용과 사용량 파악의 제약을 만듭니다. Router는 요청을 실제 작업에 맞는 모델로 보내고 하나의 base URL 또는 두 줄의 코드로 연결하며, 모델 게이트웨이는 제공자 교체와 토큰 단위 귀속을 지원합니다. 실제 작업 기준으로 같은 출력에 약 40% 낮은 비용, 초기 사용자 평균 40% 절감이 제시됐습니다.
원문 트윗 2개 보기

Veeral Patel
@vral
Monitor and control your AI spend on every provider on http:// router.com. Our early users save 40% on average. Every week, the price-intelligence-latency frontier shifts, and we expect this trend to continue. Tradeoffs between latency, reasoning, cost, service tier, open source and closed source models are shifting constantly. Router sends every request to the model that's actually best for the task and helps you control what tokens you buy. We benchmark it against real work: ~40% lower cost for the same outputs. Today we're opening it to everyone. Two lines of code or just change your base URL. No @tryramp account needed. Free through 2026, first $26 on us. Get an API key today at http:// router.com

Santiago
@svpino
Using a model gateway is the simplest way to improve the architecture of whatever you are building. • You will never be locked into a single model provider • You can swap models with a single change • Token-level attribution: you know who did what and when
Monitor and control your AI spend on every provider on https:// router.com. Our early users save 40% on average. Every week, the price-intelligence-latency frontier shifts, and we expect this trend to continue. Tradeoffs between latency, reasoning, cost, service tier, open
➖ 학생용 AI 요금제의 글로벌 확장포스트 3
Gemini와 관련 학생용 서비스가 140개 이상 국가의 대학생을 대상으로 1년 무료 이용, 학습 허브와 리서치 기능을 묶어 제공하기 시작했습니다.
- 학생이 AI 서비스의 높은 이용 한도와 저장 공간을 사용하려면 유료 요금제가 필요했던 상황에서 Google은 자격을 갖춘 대학생에게 미국에서는 Google AI Pro, 140개 이상 국가에서는 Google AI Plus를 1년 동안 무료로 제공합니다. Student Hub, study notebooks, interactive visualizations, Notebook, Flow와 Gemini Live의 Deep Research가 대상 기능으로 묶였으며, 해당 제안은 2026년 12월 31일까지 운영된다고 적시됐습니다. 학습용 기능과 이용 한도를 지역별 학생 요금제에 결합해 사용자층을 넓히는 방식입니다.
원문 트윗 2개 보기

Google Gemini
@GeminiApp
Back-to-school season is here and starting today, eligible college students can get a full year of Gemini on us: - US students: 1 year of Google AI Pro at no cost - 140+ countries: 1 year of Google AI Plus at no cost Here’s what’s new for students

Google Gemini
@GeminiApp
The student hub, study notebooks, interactive visualizations, and Deep Research in Gemini Live are rolling out today to eligible Gemini app users globally. Claim your student plan for one year at no cost. Offer runs through December 31st, 2026 for eligible students. Terms apply. Learn more here:
📈 로컬 실행과 실시간 음성·영상 모델포스트 4
오픈 모델의 로컬 실행, 실시간 음성 대화, 영상 스트리밍 LoRA가 모델 배포 위치와 생성 방식의 선택지를 넓혔습니다.
- 모델을 클라우드에서만 호출하면 데이터가 외부로 나가거나 실시간 생성에 필요한 지연시간을 맞추기 어렵습니다. S1-mini는 0.6B 오픈 가중치 언어 모델로 전사 처리를 기기 안에서 수행하고, Eleven v3 Conversational은 audio tags와 70개 이상 언어를 지원하는 실시간 음성 모델로 일반 제공에 들어갔습니다. MiniMax H3용 스트리밍 영상 LoRA는 연속 영상 생성의 가능성을 제시했지만, 실시간 생성에는 추가 추론 가속이 필요하다는 조건이 함께 제시됐습니다.
- Ornith-1.5는 9B Dense, 35B MoE, 397B MoE로 구성된 오픈소스 LLM 제품군으로 소개됐고, 게시물에서는 비교 가능한 규모의 오픈소스 모델 가운데 성능과 Claude Opus 4.8에 가까운 결과가 언급됐습니다. S1-mini의 기기 내 전사와 Ornith-1.5의 오픈소스 배포는 모델을 직접 실행하고 교체하려는 선택지를 넓히며, 음성·영상 영역에서는 LoRA와 추론 가속의 결합이 과제로 남았습니다.
원문 트윗 2개 보기
Cohere
@cohere
Locally hosted locally made
Introducing S1-mini Our first open-weights language model. A 0.6B parameter model that processes transcripts entirely on your device. Try it in app today.
ElevenLabs
@ElevenLabs
Eleven v3 Conversational, our most expressive model for realtime speech, is now generally available. For developers building voice experiences that respond with real emotion, Eleven v3 Conversational includes audio tags for fine-grained control and support across 70+ languages.
용어 해설
- Private Safety Processing
- — Zero Data Retention 환경에서 고객 데이터 통제를 유지하면서 안전 점검을 수행하는 방식입니다. 고객 인프라 안에 콘텐츠를 두고 자동화 시스템이 여러 상호작용의 위험 패턴을 찾은 뒤 제한된 안전 신호만 반환해 OpenAI 직원이 원문 프롬프트와 응답을 읽지 않도록 설계됐습니다.
- 데이터 무보존(Zero Data Retention)
- — 서비스 제공자가 고객의 프롬프트와 응답을 보존하지 않는 배포 조건입니다. 이번 맥락에서는 장기적이고 자율적인 AI 작업의 안전성을 높이면서도 민감한 콘텐츠를 고객 통제 인프라에 남겨 두려는 요구와 연결됩니다.
- 에이전트 하네스(Agent Harness)
- — AI 에이전트의 도구 호출, 컨텍스트 관리, 하위 에이전트 조정, 코드 실행을 묶어 실행하는 런타임 계층입니다. TrueForge 사례에서는 호출마다 누적 컨텍스트를 모델에 다시 보내는 비용 구조까지 관리해 모델 선택과 실행 비용에 영향을 줍니다.
- 모델 게이트웨이(Model Gateway)
- — 여러 모델 제공자에 대한 요청을 하나의 인터페이스로 연결하는 계층입니다. 요청별로 작업에 맞는 모델을 선택하고 제공자를 교체하며 토큰 사용량과 비용을 추적할 수 있어 특정 제공자에 종속되지 않는 구성을 만듭니다.
- 구조화된 출력(Structured Output)
- — 문서에서 추출한 정보를 미리 정한 스키마의 결과로 정리하는 방식입니다. Databricks의 AI Extract는 대규모 문서를 작은 작업으로 나눠 병렬 처리한 뒤 결과를 하나의 구조화된 출력으로 조정해 복잡한 중첩 스키마를 처리합니다.
- LoRA
- — 기존 모델의 전체 가중치를 다시 학습하지 않고 별도 저랭크 어댑터를 학습하는 Fine-tuning 방식입니다. 이번 포스트에서는 MiniMax H3의 영상 스트리밍 어댑터에 쓰였으며, 연속 영상 생성을 위해 추가적인 추론 가속이 필요하다는 조건이 함께 제시됐습니다.
AI 요약 · 북마크 · 개인 피드 설정 — 무료
출처 · 인용 안내
인용 시 "요약 출처: AI Trends (aitrends.kr)"를 표기하고, 사실 확인은 원문 보기 기준으로 진행해 주세요. 자세한 기준은 운영 정책을 참고해 주세요.