|
How many tokens will an old 3090 produce?
|
|
1
|
19
|
August 26, 2026
|
|
Anyone else having trouble getting a ZeroGPU allocation/reservation?
|
|
33
|
748
|
August 25, 2026
|
|
Agentic curious questions
|
|
3
|
66
|
August 25, 2026
|
|
How can I replace an object using a reference image while preserving its exact design?
|
|
1
|
18
|
August 26, 2026
|
|
Request for second username change on Hugging Face account
|
|
1
|
40
|
August 25, 2026
|
|
Building Local: My 2026 Headless AI Server Journey
|
|
8
|
337
|
August 25, 2026
|
|
TIS 2.0: Token Importance Scoring Now Eliminates Position Bias in RAG
|
|
27
|
566
|
August 25, 2026
|
|
I gave Nemotron 3.5 Lightning eyes and ears with zero training, because the architecture let me
|
|
0
|
30
|
August 24, 2026
|
|
Space stuck on "Paused" status and returning Error 503 on Factory Rebuild
|
|
3
|
67
|
August 25, 2026
|
|
Steer on a Sphere: Geometric Control of Transformer Outputs
|
|
15
|
306
|
August 25, 2026
|
|
Funasr using hugging face hub with paraformer-zh errors
|
|
1
|
47
|
August 25, 2026
|
|
Complete personal archive of Tsiolkovsky (51k sheets) — with recognition accuracy actually measured
|
|
0
|
18
|
August 25, 2026
|
|
A theoretical systems white paper outlining a four-part pipeline to eliminate Softmax denominator bloat, semantic compression loss, and hardware I/O latency in Large Language Models
|
|
28
|
400
|
August 17, 2026
|
|
Embedding models fail possibly on scientific characters no matter how the pdf is parsed
|
|
1
|
46
|
August 25, 2026
|
|
I built an experimental routing-based attention mechanism for GPT models
|
|
2
|
66
|
August 25, 2026
|
|
When is a benchmark conclusion identified across evaluator meanings?
|
|
3
|
78
|
August 24, 2026
|
|
Nemotron-3-Omni in GGUF: why audio and video need more than the usual mmproj (and where to get ones that work)
|
|
0
|
33
|
August 24, 2026
|
|
Nemotron-3-Omni GGUFs with working audio + video in llama.cpp (9 tested quants, one-pass A/V)
|
|
0
|
34
|
August 24, 2026
|
|
What Actually Matters When Comparing LLM APIs for Production?
|
|
1
|
46
|
August 24, 2026
|
|
API GPU acquisition failure — perpetually stuck at "Waiting for a GPU to become available
|
|
3
|
248
|
August 23, 2026
|
|
[Research/Code] Verified Multi-Agent Hallucination Elimination via Destructive Interference (PSAS)
|
|
4
|
98
|
August 25, 2026
|
|
Research feedback requested: Epistemic Shield — adversarial epistemic calibration for LLM outputs
|
|
2
|
57
|
August 24, 2026
|
|
Prompts what is the point?
|
|
1
|
54
|
August 24, 2026
|
|
A minimalist DSL to enforce deterministic code generation with LLMs
|
|
2
|
66
|
August 25, 2026
|
|
Same effective batch, different LoRA training time: a small TRL diagnostic
|
|
5
|
124
|
August 25, 2026
|
|
Seeking feedback on token-block context selection for long-sequence QLoRA fine-tuning
|
|
7
|
147
|
August 22, 2026
|
|
Token Cost Blew Up in Production: How Are You Handling Context Compression for Mobile-Embedded LLM Features?
|
|
0
|
52
|
August 24, 2026
|
|
Billing issue: Space not active but billing continues
|
|
7
|
564
|
August 25, 2026
|
|
Deterministic LoRA training directly through quantized GGUF serving weights
|
|
6
|
116
|
August 25, 2026
|
|
Why LLM agents keep failing (and it’s not the prompt)
|
|
6
|
445
|
August 25, 2026
|