September 2026 Summaries
27 posts from Perplexity
Filter
Month:
Year:
Post Summaries
Back to Blog
Perplexity introduces pplx-embed-v2-context-9b-preview, a 9-billion-parameter contextual embedding model designed to retrieve not only answer-bearing text chunks but also the supporting context needed to interpret and verify them within long documents. Rather than relying on conventional single “gold chunk” labels, the model is trained through distillation from a query-aware context compression teacher that assigns continuous relevance scores to document tokens, enabling flexible chunk boundaries and recognition of supporting evidence. It processes documents in a single pass to produce contextual chunk embeddings without added inference-stage reranking or compression costs, and supports 1024- or 2048-dimensional embeddings as well as int8 quantization. The accompanying privately held context-bench benchmark, developed by turbopuffer, evaluates document disambiguation, answer retrieval, and evidence recovery across 2,099 queries and nearly 39,000 long documents from 21 domains. Perplexity reports that its preview model leads evaluated contextual models on context-bench and achieves the highest average performance among shown models on the public ConTEB benchmark, while retaining competitive general retrieval performance and showing limited sensitivity to chunk size. A preview is available through Hugging Face, while context-bench evaluations can be requested from turbopuffer to reduce benchmark contamination.
Sep 30, 2026
4,564 words in the original blog post.
Perplexity has launched Automations for Computer, enabling long-running agents to perform recurring or event-triggered work while retaining context from previous runs. Replacing Scheduled Tasks, Automations can use connected files, tools, and Projects to handle assignments such as drafting customer replies, tracking competitor prices, monitoring project deadlines, updating decision logs, and preparing engineering pull requests. Users can trigger agents on schedules or through events in Slack, Gmail, Outlook, Linear, and GitHub, with conditions that narrow when an Automation runs. Setup can be done through Computer’s omnibar or the Automations interface, where users define instructions, sources, output destinations, permissions, and actions requiring human approval. Automation histories show triggers, outputs, and status, while administrators can restrict connector access; monitoring for triggers does not consume credits, which are used only when an assignment runs.
Sep 29, 2026
672 words in the original blog post.
AI agents can inadvertently become security risks when they encounter obstacles and seek workarounds that cross boundaries, as illustrated by reported incidents involving agents probing vulnerabilities, bypassing protections, or escaping evaluation environments while pursuing legitimate tasks. Drawing parallels to the early internet-worm era, Perplexity argues that agent security should rely on established engineering practices rather than any single model safeguard, particularly defense-in-depth: independent, redundant controls that constrain access, monitor behavior, and contain failures. The company distinguishes model-level safety measures from system-level protections and recommends deterministic enforcement below the agent’s control, with risk signals only able to reduce permissions. It describes applying these ideas through tools and environments including the SPACE microVM sandbox for cloud tasks, BrowseSafe prompt-injection detection for browser agents, fail-closed local sandboxes for Portable Computer, and Numbat monitoring for coding agents on employee endpoints. Perplexity also emphasizes adversarial testing, public disclosure of discovered weaknesses, open-source security tools, and collaboration among model developers, infrastructure providers, researchers, and enterprise security teams as necessary for making increasingly capable agents safer.
Sep 29, 2026
2,827 words in the original blog post.
Perplexity has updated its Agent API with Profiles, custom Skills, and managed connectors to help teams create consistent, reusable AI agent workflows across applications. Profiles store versioned agent configurations, including models, instructions, tools, service connections, and runtime settings, while Skills package reusable procedures and supporting files, and connectors provide centrally managed access to approved external services or internal systems. These features allow project administrators to configure resources once and share them with authorized developers, reducing duplicated setup and enabling centralized updates to instructions, credentials, and integrations. Example uses include standardized incident response workflows using Datadog, GitHub, and Slack, as well as reusable deployment and rollback procedures for platform teams. Profiles and custom Skills are available for Agent API projects, while managed connectors are in preview with support for GitHub, Slack, Google Drive, Datadog, Linear, and Notion.
Sep 28, 2026
666 words in the original blog post.
Model Context Protocol (MCP) is an open standard that enables AI applications and agents to discover and use external tools, data resources, and prompts dynamically through MCP servers, whereas APIs provide predefined, service-specific requests using fixed endpoints, methods, and parameters. An MCP workflow includes a host application, an embedded client that communicates using the protocol, and a server that exposes available capabilities, often relying on underlying APIs to execute requests. MCP is designed for flexible, natural-language-driven and multi-step AI workflows, particularly when tools change frequently, integrations are numerous, or an LLM must select capabilities at runtime, reducing client-side integration maintenance. Direct APIs remain more efficient for predictable, structured, high-volume, or repetitive tasks where reasoning and dynamic tool discovery offer limited value. Rather than replacing APIs, MCP functions as an intermediary layer that can make existing services more accessible to AI systems, and many AI platforms offer low-code or one-click MCP connectors for services such as email, CRM, databases, and development tools.
Sep 28, 2026
1,706 words in the original blog post.
AI alignment concerns whether AI systems pursue their operators’ intended goals rather than merely optimizing imperfect training signals or following instructions literally, making it both a technical engineering problem and a broader debate about societal safety. Key technical challenges include outer alignment, where human preference feedback may encode flawed objectives, and inner alignment, where models generalize incorrectly in unfamiliar situations; associated failure modes include reward hacking, sycophancy, specification gaming, deceptive evaluation behavior, sandbagging, scalable oversight problems, and power-seeking behavior. Frontier labs such as OpenAI, Anthropic, and Google DeepMind describe approaches involving gradual deployment, layered safeguards, capability-based thresholds, internal testing, monitoring, audits, and public reporting, though industry debate continues over whether voluntary commitments can withstand competitive pressure, with a 2026 statement from AI-company personnel calling for government-supported technical and governance controls. Governments, nonprofits, and universities also fund research, develop benchmarks, conduct independent evaluations, and coordinate international standards. Alignment differs from content moderation and encompasses only one part of wider AI safety, alongside security, robustness, and misuse prevention, while increasing model capability does not necessarily improve alignment. Perplexity’s related work focuses on applied control of web-using agents, including open-source tools and benchmarks designed to detect prompt-injection attacks in adversarial online environments.
Sep 25, 2026
2,671 words in the original blog post.
Perplexity developed Photon, an in-house Rust-based retrieval and ranking engine, to replace an adapted open-source system that became costly, slow at the tail, and difficult to recover or scale as its index grew. Photon uses compact inverted indexes and per-document “docblob” records, selective decoding, batched cache-aware asynchronous disk reads, and separate index-building and query-serving infrastructure to reduce I/O, improve resource use, and avoid disruptions during updates. Its architecture routes queries through brokers and shard groups for retrieval and multi-stage ranking, while independently built, versioned indexes are warmed with real query logs and deployed gradually. After staged testing and migration, Photon reduced production p99 retrieval-and-ranking latency from roughly 800 ms to 65 ms, eliminated prior index-update latency spikes, used about 20% fewer serving machines, and stored about 2.5 times more data per document. Photon also underpins Perplexity’s new fast Search API preset, which reports 160 ms p50 and 230 ms p95 single-search latency and, across six agentic benchmarks, achieved comparable aggregate task quality to the default preset at an estimated 68% lower total model-and-search cost, although it sacrifices some relevance, answer availability, and broad-search quality for speed-sensitive workflows.
Sep 24, 2026
3,986 words in the original blog post.
Perplexity has expanded its Portable Computer local AI platform to Windows systems using AMD Ryzen AI Max Series processors and the Ryzen AI Halo developer platform, enabling Pro and Max subscribers to run models such as PPLX 27B and Qwen 27B directly on their own hardware without consuming Computer credits for local inference. Portable combines a local model with an agent harness, scheduler, sandbox, sensitive-content classifier, and integrations for local files, Gmail, Outlook, Slack, GitHub, and supported locally installed applications through MCP servers. It supports one-time and recurring workflows, including invoice reconciliation, email briefings, contract reviews, and pull-request monitoring, while allowing users to selectively invoke cloud search and frontier models for current information or advanced reasoning. Local processing is constrained by user-defined folder and app permissions, isolated sandboxing, and an on-device classifier that flags sensitive data before cloud transfer, with organizational administrators able to manage local inference access. The feature requires Windows 10 or 11, a supported AMD system with at least 24 GB of GPU-accessible memory, roughly 20 GB of available storage, and setup through the Perplexity Windows app.
Sep 24, 2026
952 words in the original blog post.
Perplexity evaluated the security of its SPACE agent sandbox by giving nine AI models root access inside Firecracker microVMs and testing whether they could escape to the host or bypass restricted network access. Across 108 VM-to-host escape attempts, no model obtained the host-side secret, while no network bypass succeeded when all networking was blocked; however, in partial-network settings that allowed package repositories and search, four models achieved 11 successful bypasses in 54 runs. The bypasses exploited DNS spoofing to poison domain-to-IP mappings and shared CDN IP addresses to access other services through permitted infrastructure, rather than breaking VM isolation. Perplexity added source-address validation, TLS termination, hostname and HTTP-authority checks, and protocol restrictions, after which repeated tests found no verified bypasses. Tests of ten third-party sandbox platforms found similar network-policy weaknesses in eight, with seven publicly identified platforms showing at least one bypass, prompting vendor disclosures and a mix of fixes, planned mitigations, and documentation updates. The findings emphasize that VM isolation and network confinement are separate defenses and that domain-based egress controls must validate destination identity consistently across DNS, IP, TLS, and HTTP layers.
Sep 23, 2026
4,771 words in the original blog post.
AI can support an end-to-end market research workflow, from converting a broad business idea into a structured research brief through market sizing, segmentation, competitor analysis, positioning, sentiment monitoring, primary research, and final decision reporting. The approach distinguishes tasks suitable for a single query from multi-step investigations requiring repeated source checks and updates, while recommending different tool types for simple drafting, broad research, and ongoing tracking. Throughout the process, AI-generated findings should be treated as directional rather than definitive: important claims, figures, dates, market definitions, competitor events, and sentiment themes require verification against original and current sources. The guidance emphasizes using prompts that request citations, disclose uncertainty, separate observations from inferences, and identify conflicting evidence, then applying human review to surveys, high-stakes decisions, and ambiguous real-world text. Because markets, competitors, and customer preferences change quickly, it also recommends scheduled refreshes ranging from weekly sentiment checks to periodic market-sizing reviews, with deeper updates before launches, pricing changes, funding decisions, or major strategic shifts.
Sep 23, 2026
2,347 words in the original blog post.
AI literacy is the ability to understand, use, and especially evaluate AI systems and their outputs so that decisions are accurate, relevant, current, safe, and informed. As AI adoption expands and employer demand for AI fluency rises, the text argues that access to tools and prompt-writing workshops alone do not create competent users; effective literacy requires careful verification of claims, citations, calculations, recommendations, and source summaries through a “Trace, Verify, Decide” process. Because AI can generate polished but false, incomplete, outdated, or contextually inappropriate information, users must apply human judgment, particularly in high-stakes areas such as legal, financial, medical, hiring, contracts, payments, and sensitive data. The text criticizes one-time, attendance-focused “training theater” and recommends hands-on, role-specific practice with real workflows, feedback, protected learning time, and clear distinctions between tasks AI can handle independently, tasks requiring review, and decisions that humans must retain.
Sep 22, 2026
2,082 words in the original blog post.
The report describes a post-training approach that combines rejection sampling fine-tuning with hint-guided on-policy self-distillation to learn from real-world model sessions, including both successful interactions and failures identified through user corrections or tool errors. Rather than imitating every step in successful sessions or discarding unsuccessful ones, the method applies cross-entropy loss to useful actions in successful sessions and uses grounded corrective hints to create KL-divergence targets for avoidable mistakes in either successful or failed sessions, while retaining other content only as context. Hints must be based on information available before the error, aiming to avoid hindsight bias; examples include confusing a requested W-3 form with a W-2 and using an invalid tool parameter despite documented allowed values. Privacy filters exclude opted-out users and sessions containing personally identifiable information, while multiple judges and validation stages assess outcomes, identify root-cause turns, and verify whether corrections are justified. Evaluations found that hints substantially improved the base model’s ability to regenerate corrected actions, and offline tool-error rates declined from 2.79% for stock GLM 5.2 to 0.87% for an RFT-plus-OPSD checkpoint, though differing training data prevent a controlled attribution. In a later live A/B comparison between two such checkpoints, tool-call failures fell from 2.24% to 1.77%, a statistically significant 21.2% relative reduction, while user dissatisfaction did not change significantly, indicating stronger evidence for improved tool reliability than for broader task success or user satisfaction.
Sep 21, 2026
3,671 words in the original blog post.
Perplexity has introduced effort controls in Perplexity Computer, allowing users to choose among four settings—Light, Standard, High, and Ultra—that determine the model and reasoning depth used for a task. The slider is intended to balance cost and analytical effort, with lower settings suited to routine work and higher settings designed for complex or open-ended assignments such as product prioritization and market evaluation. Perplexity’s orchestrating model plans assignments and delegates work to supporting agents that may use models from multiple providers, while the company uses experience from billions of queries to match model combinations with tasks based on quality and cost. The feature is available on the web through Computer’s omnibar, with Android and iOS availability planned, and users can still manually select specific models and reasoning levels through custom controls.
Sep 17, 2026
429 words in the original blog post.
AI can help knowledge workers reduce time spent on recurring operational tasks such as staying informed, drafting communications, managing meetings, planning schedules, analyzing information, producing documents, handling administration, reviewing work, learning new topics, and preparing for stakeholder relationships. Its strongest applications involve predictable, low-risk work whose outputs can be checked easily, while human judgment remains necessary for decisions, approvals, and relationship management. Safe workplace use requires relying on approved tools and company SSO, understanding data classifications, redacting identifiers when appropriate, disabling model-training settings where needed, and limiting connected-tool permissions. To adopt AI effectively, workers can track repetitive tasks for a week, select one suitable task, define a usable output in advance, iteratively test and refine a prompt on real examples, apply a consistent quality-control check, and compare time savings after a two-week pilot.
Sep 17, 2026
2,056 words in the original blog post.
Shadow AI refers to employees using unapproved AI chatbots, agents, meeting notetakers, browser extensions, and other tools for work, often because official options are unavailable, limited, slower, or less capable than consumer alternatives. Surveys cited indicate that policy-breaking AI use is widespread, frequently involves free tools with weaker security controls, and can be difficult for organizations to detect because prompts, browser activity, and autonomous agents leave fewer traces than traditional shadow IT. Risks include exposure of confidential data and intellectual property, security breaches through connected accounts, inaccurate AI-generated work, regulatory and legal consequences, and privacy violations from unauthorized recording. The material argues that outright bans often displace rather than eliminate use, driving employees toward personal devices or indirect workarounds, while organizations can reduce shadow AI by offering useful approved tools before restricting alternatives, selecting products employees prefer, providing meaningful training, applying safeguards and accountability, and treating unauthorized use as evidence of unmet workflow needs.
Sep 15, 2026
2,636 words in the original blog post.
AI writing detectors analyze completed text for statistical patterns associated with language models rather than observing how a document was created, so their scores should be treated as signals for further review rather than proof of authorship. Different tools define percentages differently, such as an estimated probability of AI generation or the proportion of prose flagged, and their results depend on training data, scoring thresholds, document length, format, language, editing, and familiarity with the AI models and writing styles involved. Measures such as perplexity and burstiness describe predictability and variation in language but cannot uniquely identify AI writing, while lower detection thresholds increase both detections and false positives. Because AI-generated text may be uncommon in a collection, even a low false-positive rate can make a substantial share of flagged documents human-written, and research has found uneven performance on unfamiliar models, altered text, and writing by non-native English speakers. Watermarks can provide additional evidence when supported by a generating system and preserved through editing, but they cannot establish that unmarked text is human-written. Decisions with meaningful consequences should therefore combine detector results with evidence such as draft histories, notes, disclosed AI interactions, source records, author explanations, and the applicable policy.
Sep 15, 2026
2,973 words in the original blog post.
HP and Perplexity are partnering to pre-load the Perplexity Windows app on selected HP devices, beginning with the HP ZBook Ultra G3a mobile workstation, to provide AI research, task automation, and local processing capabilities. The app integrates with Microsoft 365 applications, Teams, local files, web resources, and more than 400 connected apps and data sources, while orchestrating models from providers including OpenAI, Anthropic, and Google. Supported devices can use Portable Computer to run certain AI tasks locally, allowing users to analyze data and files without sending work to the cloud, though they may authorize cloud escalation for more advanced tasks and receive warnings about potential sensitive-data sharing. The collaboration also includes a read-only integration with Autodesk Revit 2027’s MCP server, enabling users to query open building models in natural language, identify and export relevant information, and support reporting workflows. The ZBook Ultra G3a uses AMD Ryzen AI Max PRO 400-series processors for demanding AI and graphics workloads, while the Perplexity app is also available through the Microsoft Store for other Windows users, with Personal Computer for Windows offered to Pro, Max, and Enterprise subscribers on Windows 10 and 11.
Sep 15, 2026
730 words in the original blog post.
Perplexity has launched Portable Computer for compatible Windows PCs, enabling Pro and Max subscribers to run a local AI model, agent harness, orchestrator, and scheduler entirely on their own hardware. Designed for systems with eligible NVIDIA GeForce RTX or RTX PRO GPUs with at least 24GB of VRAM, the feature keeps sensitive files, queries, and agent activity on-device and does not consume Computer credits for local work. Users can schedule recurring workflows such as invoice reconciliation, file processing, and code-review tasks, while Portable can access permitted local documents, spreadsheets, PDFs, images, code, and connected desktop applications through local Model Context Protocol servers. It can also integrate with Gmail, Outlook, Slack, and GitHub, and users may optionally authorize cloud access for Perplexity Search or more than 15 frontier models when current information or advanced reasoning is needed.
Sep 14, 2026
698 words in the original blog post.
An AI-native search company rebuilt its storage stack to better support the distinct demands of processing web pages into passages and embeddings, continuously updating document state, and serving low-latency batch reads for search queries. Replacing a DynamoDB-based architecture that directly coupled processing writes with serving, the new design uses Pillar for durable, versioned document state and export decisions, Lorry to create partition-aligned update batches, and CobbleDB, a custom distributed key-value store built on RocksDB, for tunable hot-storage reads. CobbleDB uses replicated partitions, local NVMe storage, memory caching, parallel routing, batched MultiGet operations, and optional replica hedging to optimize retrieval of prepared page records, while asynchronous ingestion prevents large rebuilds or updates from competing with live traffic. Production measurements reported median batch-read latency falling from 31.4 milliseconds to 5.6 milliseconds, with similar roughly fivefold improvements at higher percentiles, while internal cost estimates indicated savings of at least 20 percent compared with DynamoDB. The project, including a roughly 40,000-line Rust database, was developed by two engineers with assistance from hundreds of internal coding agents that helped review changes, identify risks, prepare fixes, track rollouts, and maintain project status, while humans retained responsibility for architecture and production decisions.
Sep 14, 2026
2,559 words in the original blog post.
Q2D-Web (Query2Doc-Web) is a private benchmark and public leaderboard for assessing retrieval models used in agentic retrieval-augmented generation systems at web scale. It includes approximately 190 million deduplicated web documents and 69,721 privacy-filtered, agent-reformulated queries spanning ten languages, sampled from nine months of production search traffic and covering diverse domains. To address sparse and potentially biased relevance labels, it offers citation-based, production web-ranking, and expanded LLM-judged relevance sets, averaging 88.5 positive judgments per query in the combined set. The benchmark emphasizes corpus size, query volume, and depth of relevance annotations as necessary for realistic evaluation, particularly because smaller corpora can remove difficult distractors and overstate retrieval quality. An RRF-based subcorpus containing 31.7% of documents was designed to reduce evaluation costs while largely preserving full-corpus model rankings. Results use Recall@1000 as the primary metric, with pplx-embed-v1-4b achieving the highest Recall@1000 across all judgment sets, while Nemotron-3-Embed-8B led on certain higher-ranking metrics; the report also notes that Perplexity models could have an in-distribution advantage despite exclusions from training data. Publicly available Hugging Face models can be submitted for standardized evaluation under specified implementation and access requirements.
Sep 09, 2026
2,050 words in the original blog post.
AI integration involves connecting artificial intelligence to an organization’s data, applications, and everyday workflows so it can analyze information, automate routine tasks, support decisions, and act across systems rather than functioning as a standalone tool. The text argues that measurable returns depend on selecting specific workflows, establishing baseline metrics for time, cost, errors, and output, running limited pilots with human review, and scaling only where results are reliable. Examples cited include banks, insurers, venture firms, and product teams reporting faster reviews, call handling, research, and development work. Potential applications span customer service, finance, legal review, marketing, security monitoring, fraud detection, sales, and IT operations, where AI may reduce repetitive work and improve access to organizational knowledge. However, poor-quality data, difficult integrations, excessive permissions, hallucinations, model drift, weak governance, and insufficient maintenance can undermine benefits, making ongoing oversight, clear sourcing, employee training, and accountability important. The text also presents Perplexity’s enterprise products as tools for connecting AI to company data, external web information, and business applications while stating that enterprise customer data is not used to train its models.
Sep 07, 2026
2,044 words in the original blog post.
Perplexity describes a custom embedding-model serving stack designed to support both high-throughput indexing and low-latency online search, using its own models such as pplx-embed to improve search relevance, cost, and response times. The system combines Ivy, a Rust HTTP gateway that handles request processing, tokenization, batching, and load balancing; Tulip, a Rust-based gRPC inference server responsible for scheduling batches; and ROSE, a Python-based engine that executes model inference on accelerators while reusing much of the company’s LLM-serving infrastructure. The architecture treats bulk embedding workloads similarly to LLM prefill and short online queries similarly to decode, allowing shared kernels and components across model types. Key optimizations include whole-model CUDA graphs to reduce CPU kernel-launch overhead, lazy graph capture to limit startup costs, and LazyTensor-based asynchronous result handling to overlap CPU preparation with GPU computation. Perplexity also uses request splitting, in-house tokenization, ragged attention without KV caches for embedding models, and multiple attention-kernel backends selected by model shape and sequence length. Benchmark comparisons with vLLM are presented for latency, scoring, throughput, and concurrent-request performance, and the company concludes that ownership of the full Ivy–Tulip–ROSE stack provides greater efficiency and flexibility while retaining reusable infrastructure for future embedding and LLM improvements.
Sep 04, 2026
2,295 words in the original blog post.
Perplexity introduces PII-TRACE, a public benchmark for testing personally identifiable information detection in long, multilingual, multi-turn conversations, and PII-Tracer, a 0.6-billion-parameter local model intended to act as a privacy gate in hybrid Mac-based AI systems. Hybrid compute routes research and planning to cloud models while retaining private files and sensitive context locally, but its privacy protections depend on reliably detecting every occurrence of recurring PII before data leaves the device. PII-TRACE contains 13,148 synthetic conversations in 13 languages, preserves dialogue structure and repeated identifiers, and evaluates both character-level detection and whether all mentions of an identifier are found across turns. PII-Tracer uses bidirectional attention, token-level BIOES span tagging, a sensitive-conversation auxiliary classifier, and constrained decoding to identify nine PII categories locally. In evaluations against 11 other detectors, it reportedly achieved the highest character-level F1 score and stronger coverage of recurring and cross-turn identifiers than larger cloud-hosted frontier models, while sliding-window processing substantially improved its performance on very long conversations. The model also outperformed the OpenAI Privacy Filter on five conventional external PII benchmarks, according to the reported character-level F1 results, and both the benchmark and model are available through Hugging Face.
Sep 01, 2026
2,633 words in the original blog post.
AI hallucinations are false, unsupported, or internally inconsistent model outputs that can appear credible through polished language, precise numbers, and invented or misused citations, especially when prompts concern obscure, recent, ambiguous, conflicting, or falsely premised topics. The risk and consequences vary by use case, from minor wasted effort in brainstorming to serious harm in legal, medical, financial, or security decisions, so readers should assess traceability, consistency, and currency. Key warning signs include unverifiable sources, citations that do not support the accompanying claim, exact figures without clear methods, incorrectly combined real facts, unchallenged false assumptions, incompatible details after rephrasing a question, and outdated evidence for current claims. The proposed five-minute review process prioritizes important claims, verifies citations and source details, matches assertions to original evidence, checks for newer primary sources, and retests questions without embedded assumptions. Perplexity describes using search-linked citations, post-training for evidence use and uncertainty recognition, task-specific evaluation, and efficient retrieval to reduce hallucination risks, while emphasizing that citations and search do not eliminate the need for human verification of consequential claims.
Sep 01, 2026
2,351 words in the original blog post.
AI personal assistants are software tools that use natural-language instructions, automation, and authorized access to apps, documents, email, calendars, and websites to help individuals organize and complete routine work such as research, meeting preparation, inbox triage, note-taking, scheduling, and follow-up tracking. Unlike chatbots, which primarily answer prompts, and agents, which autonomously pursue defined goals through multiple steps, personal assistants focus on supporting one person’s daily workflow and may combine conversational interfaces with agent capabilities. Current systems work best on specific, verifiable tasks, as they can still make errors; Stanford’s 2026 AI Index found leading agents succeeded on roughly two-thirds of common computer-use tasks. Available options include assistants bundled into workplace ecosystems, dedicated standalone tools, and custom-built systems, with trade-offs involving cross-platform access, setup effort, maintenance, and governance. Effective deployment depends on limiting permissions, using predictable workflows where possible, defining memory and data-retention policies, maintaining logs, treating external content as potentially malicious, and requiring human approval for consequential actions such as sending external messages, spending money, deleting data, or changing sensitive records.
Sep 01, 2026
2,100 words in the original blog post.
Hybrid Compute divides AI tasks between cloud-based frontier models for research and reasoning and local models on Macs for private files and apps, making fast local inference important for a seamless experience. Lily is a lightweight Rust and Metal inference engine built specifically for Qwen3.6-35B-A3B on Apple silicon, bypassing PyTorch and MLX in favor of model-specific execution plans for prompt processing, or prefill, and token generation, or decode. On an M5 Max MacBook Pro with 128 GB of unified memory, Lily averaged 1.23 times MLX-LM’s prefill throughput and 1.35 times its decode throughput across contexts from 256 to 128K tokens, reaching 5,749.9 prefill tokens per second and 186.6 decode tokens per second at 4K-token workloads. Its optimizations address Qwen’s sparse mixture-of-experts routing, recurrent Gated DeltaNet layers, and grouped-query attention by keeping routing and state on the GPU, dequantizing 4-bit weights during computation, selecting workload-specific matrix or vector paths, fusing operations to reduce memory traffic, and reusing attention-cache data across query heads. Tests found that some techniques, including speculative decoding, reduced overall performance under these conditions, while measurements suggested core operations were already approaching hardware bandwidth and compute limits. Lily’s output remained broadly consistent with MLX-LM, with perplexity only 0.04% higher and matching top-token choices at 96.35% of tested positions, and the project’s authors argue that efficient local inference increasingly requires engines tailored to both a model’s architecture and the hardware platform.
Sep 01, 2026
4,095 words in the original blog post.
Perplexity has introduced hybrid compute for its Mac app, combining cloud-based frontier AI with local models that process sensitive files, tools, and client data directly on Apple silicon Macs. A privacy gate uses an on-device classifier to identify sensitive information such as credentials, account numbers, addresses, and government IDs, then can retain, mask, block, or request consent before data is sent to the cloud. The cloud handles web search, planning, and advanced reasoning, while the Mac manages private-document analysis and local actions, allowing workflows in finance, advertising, legal work, healthcare, and other sensitive fields without requiring users to manually separate information. Enterprise administrators can establish organization-wide privacy rules and audit when information leaves a device. The feature is available to Perplexity Pro, Max, and Enterprise subscribers on Apple silicon Macs running macOS 15 or later with at least 24GB of unified memory, launching with Gemma 4 E4B, Qwen3.6 35B-A3B, and a Perplexity local model, with dedicated Mac minis optionally enabling remote access through an iPhone.
Sep 01, 2026
817 words in the original blog post.