Dnotitia Inc. (Dnotitia), an AI data infrastructure and semiconductor company, today announced that the first ASIC samples of its Vector Data Processing Unit (VDPU) have returned from fabrication, with chip-level characterization now underway. At AI Infra Summit 2026, held Sept. 15-17 at the Santa Clara Convention Center, the company showcased a server-scale VDPU architecture designed for vector retrieval workloads in retrieval-augmented generation (RAG) and agentic AI.
The showcase followed the company’s first public display of its VDPU chip and accelerator card at the Future of Memory and Storage (FMS) 2026, where Dnotitia received a Best of Show Award, and marked the next step from chip-level demonstration to server-scale evaluation.
On Dnotitia’s FPGA evaluation platform, a server equipped with four VDPU cards delivered up to 5.77x the vector-search throughput of the same software stack running on a dual-socket CPU-only server, while maintaining equal or better recall. In a 4,096-dimensional multimodal workload, VDPU reduced host CPU use during index building by 92% and host memory by 73%, freeing host CPU resources for applications. All performance figures were measured on the FPGA platform and do not represent final ASIC performance.
“Agentic AI is shifting the AI infrastructure bottleneck from model compute toward retrieval,” said Se-Hyun Yang, Chief Technology Officer of Dnotitia. “As models search and verify information repeatedly, retrieval needs its own processing layer. VDPU is designed to give CPU capacity back to applications and keep GPU HBM focused on model execution. At AI Infra Summit, we showcased VDPU at server scale and opened discussions with infrastructure partners around evaluation and integration.”
The FPGA platform has been validated with FAISS, Milvus and hnswlib across brute-force KNN, IVF, NSW and HNSW indexes. Beyond the stacks already ported to the FPGA platform, Dnotitia plans to support a broader range of vector libraries and databases on the ASIC so that VDPU can work with the environments customers already run.
Dnotitia’s first-generation VDPU ASIC is currently undergoing chip-level characterization. The company plans to begin ASIC-based VDPU evaluations in Q4 2026. Dnotitia is targeting up to 10x vector-search performance versus a CPU-based server with its VDPU ASIC-based server.
During the summit, Se-Hyun Yang presented “Rethinking AI Infrastructure with Dedicated Vector Silicon,” outlining the architecture and performance results behind VDPU. Dnotitia also met with prospective customers and infrastructure partners at Booth 205 to discuss VDPU evaluations, proof-of-concept projects and potential integration into existing server, storage, and AI infrastructure.
Following the event, the company is continuing discussions with server, storage, memory and semiconductor companies, as well as vector database providers and AI framework developers, as it expands server-scale VDPU evaluations and prepares for ASIC-based evaluation in Q4 2026.
Facts Only
* Dnotitia announced the return of first ASIC samples for its Vector Data Processing Unit (VDPU) from fabrication.
* Chip-level characterization of the VDPU is currently underway.
* At AI Infra Summit 2026, Dnotitia showcased a server-scale VDPU architecture for vector retrieval workloads in RAG and agentic AI.
* On an FPGA evaluation platform, a server with four VDPU cards achieved up to 5.77x vector-search throughput over a dual-socket CPU-only server, maintaining equal or better recall.
* In a 4,096-dimensional multimodal workload, the VDPU reduced host CPU use during index building by 92% and host memory by 73%.
* Performance figures were measured on the FPGA platform and do not reflect final ASIC performance.
* The FPGA platform was validated with FAISS, Milvus, and hnswlib across various indexes (brute-force KNN, IVF, NSW, HNSW).
* Dnotitia plans to support a broader range of vector libraries and databases on the ASIC.
* The company plans to begin ASIC-based VDPU evaluations in Q4 2026.
Executive Summary
Full Take
The narrative positions dedicated vector silicon as a necessary architectural response to the growing demands of agentic AI, framing retrieval as the current bottleneck rather than model computation. This suggests a shift in infrastructure priorities where the efficiency of information retrieval dictates overall system performance, moving processing away from general-purpose CPU/GPU compute toward specialized vector operations. The demonstration on FPGA platforms provides strong empirical evidence—the 5.77x throughput and significant reduction in host resource utilization during indexing—suggesting that decoupling retrieval from host resources yields substantial gains for specific workloads like RAG.
The pattern observed is the movement from abstract capability presentation (showcasing architecture) to concrete, measurable performance benchmarks (FPGA evaluation). The context provided by validating against established vector libraries (FAISS, Milvus) establishes a critical bridge between theoretical design and practical utility. However, the transition remains gated by chip-level characterization for the final ASIC performance claims, creating an inherent tension between immediate, platform-based success and future hardware realization. A key implication is the potential restructuring of the AI infrastructure stack, where specialized silicon could fundamentally alter where compute resources are allocated in complex AI systems. The missing piece lies in assessing the long-term integration viability and cost implications of this specialization versus generalized parallel processing approaches when deployed at massive scale outside controlled FPGA environments.
BRIDGE QUESTIONS: If the ASIC evaluations target a 10x vector-search performance over a CPU server, what specific constraints—power, latency consistency, or memory bandwidth—will govern the practical deployment of this silicon in heterogeneous systems? How will Dnotitia ensure that expanding library support on the ASIC maintains the same level of abstraction and interoperability achieved across external software stacks? What metrics beyond throughput and host reduction are necessary to fully evaluate the long-term value proposition for infrastructure partners integrating these specialized units?
