This is the English edition. 한국어판 and 日本語版 are also available.

Perplexity Portable Computer: The 24GB-VRAM Local AI Agent Explained

2026-09-26 · AI · United States · Zoogom Editorial

#Perplexity#Portable Computer#local AI#AI agents#RTX#privacy

A local AI workstation processing development and document tasks on its own hardware

Perplexity Portable Computer is not merely a chatbot model running on a desktop. It places the agent harness, orchestrator, planner, local model, and sandbox on a supported Windows or Linux machine. When a step needs the web or a frontier cloud model, the system can ask for approval, send that step outside the computer, and return the result to the local run.

That design treats local and cloud processing as parts of one workflow rather than mutually exclusive products. A repository-scale migration can stay on the user’s GPU while one current-documentation lookup goes to an external service. The distinction is valuable only if the approval request makes the outgoing data and destination understandable.

Perplexity also says work completed by local models has no per-token charge. That does not make the system free: subscription, GPU hardware, electricity, administration, and human review remain. Nor does “local-first” mean “always offline.” Cloud escalation, search, and connected apps can move information beyond the device after approval.

Three things to know

  1. The main product is an orchestration boundary: local execution by default, optional cloud escalation for selected steps.
  2. Hardware requirements are substantial—supported Windows or Linux hardware with at least 24GB of VRAM—and macOS is not on the current list.
  3. A responsible evaluation must cover permissions, sandbox limits, audit receipts, power use, and recovery from bad actions, not only model quality.

What runs where

What runs where: Component, Default location, Role, What to verify

The official Perplexity Portable Computer page says the local stack includes the harness, orchestrator, planner, and PPLX 27B or an available Qwen model. A step that needs the web or higher-end reasoning can be routed, after approval, to one of more than 15 cloud models.

The meaningful unit is the step, not necessarily the entire conversation. That could reduce exposure compared with uploading a complete workspace to a cloud agent. But the system must show whether it sends a question alone, attached file excerpts, prior context, tool output, or all of the above. A generic “allow” button is not enough for sensitive work.

Local file and code processing followed by an approval gate for selected cloud steps

Hardware and model availability

Perplexity lists NVIDIA DGX Spark, a supported NVIDIA RTX GPU, or AMD Ryzen AI Max hardware with at least 24GB of VRAM. The requirement is GPU-addressable memory on a supported system, not simply 24GB of ordinary system RAM.

On DGX Spark, the published local choices are PPLX 27B and Qwen 3.8 27B. Windows supports local setup on listed RTX or Ryzen AI Max systems. Linux RTX PCs currently get PPLX 27B; the page explicitly says Qwen 3.8 27B is not available on Linux RTX PCs. Only one local model runs at a time. NVIDIA Nemotron 3.5 Lightning is described as coming soon and should not be treated as a shipping capability.

Buyers should consider sustained operation, not only a short benchmark. Long agent jobs can occupy a GPU for hours. Cooling, fan noise, driver stability, sleep and resume behavior, model storage, and remote-session reliability may matter more than peak tokens per second.

What “no per-token charge” does and does not mean

Local inference avoids metering every local token through an external model API. That can make repeated repository transformations, large file summaries, or batch classification more predictable. The total cost still includes:

An organization with suitable hardware and high local throughput may see a clear advantage. A user who asks occasional short questions and must buy a new workstation may find a cloud-only subscription simpler and less expensive.

Connectors expand both utility and risk

The product page names Gmail, Outlook, Slack, and GitHub connectors. In one run, an agent could read an inbox, inspect a repository, prepare a summary, and post it to a channel. Each added action expands the authorization boundary.

Start with read-only accounts, disposable repositories, and reversible operations. Preserve per-action confirmation for sending messages, pushing branches, deleting files, changing schedules, or modifying accounts. Check whether a one-time grant becomes a durable permission and whether scheduled tasks inherit the same authority when no user is present.

A local sandbox is useful but not self-defining. Administrators still need to inspect allowed folders, outbound network destinations, shell capabilities, access to environment variables, child processes, and connector tokens. Source repositories and important documents should have independent backups. A human should review a diff before a large generated change is committed or sent.

A practical U.S. evaluation plan

Choose three representative, non-sensitive jobs: categorize a large local file set, explain an unfamiliar codebase, and diagnose a failing test suite. Record elapsed time, VRAM use, power consumption, model errors, and human review time. Quality per finished task is more useful than raw throughput.

Run the same jobs with all cloud escalation denied. Then enable only web search, followed by one frontier model, to identify the benefit and data cost of each route. Test connectors last, beginning with a read-only GitHub repository before mail or Slack posting.

For corporate deployments, map each route to data-classification policy. Source code, customer records, protected health information, student records, export-controlled material, and attorney-client communications may require different controls. A user approval does not override contractual, regulatory, or employer restrictions.

Who should consider it now

Portable Computer is most compelling for developers and organizations with large local code or document workloads, qualified GPU hardware, and a desire to control which steps leave the machine. It is a weaker fit for Mac-only users, thin-client fleets, teams without workstation administration, or environments that require every task to remain offline.

The product’s important idea is not that local models replace the cloud. It is that an agent can make compute location and data movement explicit at each stage. Success will depend on whether approvals are specific, failures stop safely, receipts show what left the device, and local results are good enough to justify the hardware.

Sources and use notice

This article independently analyzes Perplexity’s published capabilities and requirements. Performance, cost, and security benefits require independent testing in each environment. It does not reproduce official product screens, logos, or third-party copyrighted media.

Source: Perplexity · Includes original screenshots or graphics