Building FAR AI | Cheaper, faster and scalable AI inference | Based on distributed compute | Powered by @Dizzaract farlabs.ai/

Be First to Try FAR AI 👉🏼
Pinned Tweet
Your GPU could do more than sit idle. Node registrations for FAR AI are now open. See your estimated output, submit your details, and secure your place early in the network. Register here: farlabs.ai/join-network#beco…
14
90
287
195,606
Which metric tells you the most about real inference performance?
43% Time to first token
0% End-to-end latency
29% Successful completions
29% Performance consistency
7 votes • Final results
6
5
4,080
As AI demand grows, a single cluster may not always offer the right capacity, latency or geographic proximity for every workload. The ability to determine where each inference request should run and route it accordingly, is becoming a critical infrastructure layer of its own. This shifts the challenge from simply adding compute to coordinating available compute more intelligently. FAR AI is built around this model: bringing independent GPU operators into a shared inference network and routing workloads based on real-time infrastructure conditions.
5
3
15
3,814
What will matter most for scaling AI infrastructure by 2030?
48% Power availability
9% Cooling capacity
30% More GPU capacity
13% Smarter use of compute
23 votes • Final results
1
8
4,740
A study published in Joule estimates that a typical frontier-model query consumes a median 0.31 Wh of energy. When queries are 15× longer in a test-time scaling scenario, estimated consumption rises to 3.91 Wh. This shows why inference workloads cannot be understood through one total latency metric. Two requests to the same model can place very different demands on infrastructure depending on how long the model spends reasoning before generating a response. FAR AI separates reasoning time from response generation time and records energy consumption for completed requests, giving teams clearer visibility into the performance and energy profile of each workload.
1
1
11
4,127
Strong hardware alone does not make a node reliable. Uptime, successful job completion, end-to-end latency and previous incidents all affect how well the network performs. FAR AI’s Reliability Score converts verified behavior into a rolling performance record. The orchestrator routes the request to the highest-scoring nodes. Persistent anomalies lower a node’s routing priority, while serious violations remove it from routing until resolved. In FAR AI’s architecture, trust is earned through verified performance over time.
10
4
26
3,922
Gartner expects global AI inference spending to reach $23.3 billion, ahead of the $19 billion allocated to training. Training develops model capabilities. Inference puts those capabilities to work in live applications, where every request creates an operational workload. As production demand scales, cost efficiency, latency and reliability become critical. This is what FAR AI is built to handle through distributed inference and designed for lower costs, lower latency and greater reliability. AI builders, join early access: farlabs.ai/join-as-ai-builde…
8
1
17
3,648
CNCF’s AI conformance program saw the number of certified platforms grow by nearly 70% in just a few months, reflecting growing demand for infrastructure built around common interfaces. FAR AI brings that simplicity to distributed inference through a single, consistent application interface. Behind it, the orchestrator routes each request to nodes with the required model loaded and sufficient hardware, so applications do not need to adapt to individual GPUs or hardware tiers. AI builders, join early access: farlabs.ai/join-as-ai-builde…
5
17
4,162
Agentic AI benefits from inference closer to where data and actions happen. But larger open-weight models can require more compute than a single local machine can provide. FAR AI connects compatible GPUs across a distributed network and supports multi-machine inference through a standard API, while managing deployment and request coordination. Building with AI? Register for early access: farlabs.ai/join-as-ai-builde…
16
6
39
18,492
Google Cloud research found that 83% of organizations need infrastructure upgrades to support production-grade AI agents. A single agent request can initiate long reasoning loops, tool calls, database queries and multiple downstream actions, creating workloads that are increasingly difficult to predict and manage. FAR AI gives organizations access to coordinated, distributed inference without requiring them to manage the underlying serving infrastructure. Latency, cache and energy metrics provide visibility into workload performance across the network.
2
2
17
16,274
Where do you think the biggest compute opportunity lies?
14% Idle data-centre capacity
23% Enterprise infrastructure
27% Independent GPU providers
36% A mix of all three
22 votes • Final results
3
10
4,207
Building with AI? Register for early access to FAR AI and claim 1M free credits when access opens 👇 farlabs.ai/join-as-ai-builde…
AI inference is becoming one of the biggest recurring costs for builders. Even though the cost per token has fallen dramatically, AI usage is growing even faster. By 2030, inference is projected to account for 37% of global data center workloads, making it one of the largest infrastructure challenges in AI. At the same time, there's over 100 gigawatts of idle compute sitting unused around the world. FAR AI unlocks that capacity to deliver lower-cost inference, with reliable execution, secure and private workloads and intelligent routing for production AI applications. Register for Early Access: farlabs.ai/join-as-ai-builde…
3
4
26
29,095
Boris Cherny, creator of Claude Code, says agent loops now produce around 30% of his code on an average day. These loops can review code, run tests, track feedback and continue working in the background. The result is not one inference request, but a chain of model calls that may continue for hours. Speed alone is not enough. If a node becomes unavailable during the workflow, later steps may be delayed or interrupted. FAR AI’s Reliability Score evaluates uptime, job completion, latency and incident history. The Orchestrator uses this performance record when routing requests, favoring nodes that have demonstrated greater consistency over time. 👇Register as a builder for early access to FAR AI: farlabs.ai/join-as-ai-builde…
1
2
18
19,222
A June 2026 Carnegie Endowment report, citing an IEA estimate, found that reducing data-center grid demand just 1% of the time could unlock around 110 GW of additional capacity across the US and EU. That could mean scheduling flexible workloads outside peak periods or moving them to regions where energy demand is lower. 👇
4
3
15
22,054
FAR AI uses existing GPUs across a distributed network instead of tying every new inference workload to additional centralized infrastructure. This distributed model creates the foundation for more flexible compute and better use of available energy infrastructure. Put your available GPU capacity to work. Join the waitlist to try FAR AI as a node operator: farlabs.ai/join-network#beco…
1
5
440
Which shift will have the greatest impact on AI infrastructure?
40% Growth of AI agents
10% Open-weight adoption
29% Distributed compute
21% Specialised hardware
48 votes • Final results
1
2
11
20,921
Open-weight models accounted for 29% of token volume on Vercel AI Gateway in June, up from 11% in April, while representing less than 4% of spend. Roughly 1 in 8 enterprise customers now run an open-weight model in production.
2
2
14
20,382
Their growing use reflects what developers are looking for: more choice, flexibility and better economics for high-volume workloads. FAR AI makes these models easier to access through a single API. When a model is requested, the network matches it with suitable hardware and nodes with stronger reliability records, without developers having to operate the underlying GPU infrastructure themselves.
6
535
Vista Equity Partners’ research, informed by production agents across 50+ portfolio companies, found that inference costs could be reduced by more than 80%, with accuracy staying within 1–2% of the most expensive alternative. The difference comes down to smarter model selection, infrastructure and agent design. FAR AI brings this approach to distributed infrastructure, coordinating open models and GPU capacity so workloads can run on resources suited to their requirements.
5
4
22
18,305
Where do you fit into the compute ecosystem today?
17% Building and need compute
32% Have GPU capacity to offe
22% A bit of both
29% Just exploring
41 votes • Final results
5
1
10
20,040