Titbits about Teradata Enterprise Vector Store

—

by

in

Teradata Enterprise Vector Store (EVS) is Teradata’s managed vector retrieval platform for Retrieval Augmented Generation (RAG), AI agents, and enterprise search workloads on governed data.

  • Keeps data in vector format,
  • Searches high-dimensional vector embeddings efficiently,
  • Provides database functionality beyond vectors,
  • Allows integration of vector and relational data,
  • Supplies complex query support,
  • Admits flexible data models, and
  • Includes advanced indexing and optimisation.
Where Teradata Enterprise Vector Store Sits

Vectors are an additional data type that Teradata’s parallel architecture can manage and process.

Additionally, Teradata Enterprise Vector Store is accessible via APIs, which open new opportunities for users to design applications and processes, such as building their own Natural Language applications using Teradata as the foundational compute engine and storage (in vector format).

Note that Teradata Enterprise Vector Store works in two possible ways:

  • As a stand-alone service, or
  • As part of Teradata AI Agentic Platform, along with AI Studio.
Teradata Agentic Platform

Enterprise Vector Store is part of the Teradata Database engine since version 20.0.

Main Features

  • Vector data storage: It enables storage of large datasets of vector representations, allowing similarity comparisons based on vector distances.
  • Scalability for large datasets: Designed to handle massive amounts of vector data, leveraging Teradata’s established capabilities for large-scale data warehousing.
  • Integration with AI workflows: You can easily integrate your machine learning models (including LLMs) with the vector store to perform similarity searches, generate recommendations, or analyse relationships within complex data.
  • Enterprise-grade features: Teradata emphasises the security, governance, and reliability aspects of its vector store, making it suitable for critical business applications.

Algorithms Available in Teradata Enterprise Vector Store

This section lists algorithms at the time of publishing this post. See the online documentation in the Documentation for Teradata Enterprise Vector Store section for up-to-date information on all algorithms available in EVS.

  • Indexing and search algorithms: FLAT, IVF_FLAT, HNSW.
  • Embedding generation: Teradata database SQLMR function AI_TextEmbedding and BYOM function ONNXEmbedding.
  • Search and response generation
    • The Semantic Search uses the Teradata SQLMR predict functions.
    • Hybrid Search combines semantic search (dense vector similarity) with keyword-based lexical search (BM25).
  • Supported third-party AI models on EVS: Third-party models are available through the Teradata account on provisioned systems without providing credentials. Teradata supports models bundled in AWS Bedrock, Azure OpenAI and Google Cloud (including Vertex AI), provided they are available in the region where the Teradata account is.
  • Custom models from the user account: You can bring your own model subscriptions using API keys and environment credentials or leverage open-source models. For example, you can bring your cloud accounts (AWS, Azure, GCP) to use in EVS deployed on the Teradata Cloud account. Teradata does not store user account credentials; users must specify them in each API call that requires model access.

How to Use Enterprise Vector Store

You can perform operations on Teradata Enterprise Vector Store through the following means:

  • Python via the teradatagenai package, a Teradata-developed Generative AI package.
    • You must install this package first.
  • Teradata in-database SQLMR framework through the SQL Interface.
    • These functions let you perform embedding, indexing, and search in a massive parallel, scalable, and concurrent way.
    • The SQL Interface uses the Teradata stored procedures created in the TD_SYSAI database.
    • The SQL Interface has been available since VantageCloud Enterprise 3.4.
  • REST APIs.
    • Open APIs allow you to build your own Natural Language applications, including Retrieval-Augmented Generation (RAG), using Teradata as the foundational compute and storage engine.
  • User Interface — Hosted on AI Studio.
    • AI Studio is available on certain versions of Teradata Cloud and Factory.

In other words, you can access the Teradata Enterprise Vector Store through:

  • The SQL Interface in the Teradata SQL Engine.
  • The Vector Store UI available in AI Studio.
  • Other platforms require API or Python SDK access.

Data Ingestion in the Teradata Enterprise Vector Store

Data ingestion in EVS uses the mechanisms described in this section. Note that when paired with AI Studio in the Teradata Agentic Platform, EVS does not use the MCP Server (hosted in the AI Studio cluster).

Structured data

Vector collections can ingest data from Teradata database tables and views. These tables contain text data that users need to search for:

  • Similarity,
  • Metadata associated with the text data for context, and
  • Optionally, pre-generated embeddings for the text data. When embeddings are not provided, the specified text data columns are vectorised using an embedding model and integrated into the vector collection.

Examples include articles, abstracts, product reviews, customer complaints, chunked documents, etc.

Unstructured data

Vector collections can natively ingest data from PDF, Excel, CSV, and Parquet files. i.e. through the Enterprise Vector Store UI. Users can also ingest them through Unstructured (unstructured.io) and nv-ingest. Other file types, such as JSON, require Unstructured or nv-ingest.

In any case, files must be in:

  • The local system, or
  • Remote object storage (AWS S3, Azure Blob Storage, Google Cloud Storage).
    • EVS downloads the files from these remote locations and processes them in batches.
    • Authentication and connectivity to access the files are similar to what users do for NOS processing.

PDF Ingestion

Internal ingestion

The EVS maintains an internal PDF ingest pipeline for parsing and chunking. There are three chunking options available:

  • Basic: The PDF document is split into fixed-length chunks. It extracts only text data from PDFs; images, charts, and tables are ignored.
  • Unstructured: The PDF document is parsed internally using the open-source unstructured.io library. The files are parsed based on document structure, and chunks are dynamically created based on document layout. Tables, images, and charts can also be extracted; however, scalability is limited. For a fully scalable solution, use the Unstructured OEM connector (from unstructured.io).
  • nv-ingest: The PDF file is parsed using the nv-ingest service, available on deployments with NVIDIA nv-ingest capability. The files are sent to nv-ingest for parsing and chunking. The service can also extract tables, images, and charts.
External ingestion

You can ingest files externally into EVS using the Unstructured OEM connector or nv-ingest. This is the preferred approach for large-scale document ingestion.

Both Unstructured and nv-ingest pipelines use specialised microservices to find, contextualise, and extract text, tables, charts, and images for use in downstream applications.

You can create a vector collection from the resulting JSON either by supplying JSON directly to the vector collection’s ingest API or by extracting fields from the JSON into a table and creating a vector collection from that table.

JSON

Teradata Vector Store accepts JSON as input. The JSON file can have any structure and may come from Unstructured, nv-ingest, or another source. JSON files are expected to have text data in chunked form. They can also include embedding elements and other metadata fields. The Vector Store Ingest API provides schema options to customise the fields extracted from the JSONs.

CSV

Teradata supports CSV input files. CSV files should include a header row, followed by data, and may include text, metadata, and embedding columns. Users can provide a custom schema to extract selected fields from CSV files.

Parquet

Teradata also supports input Parquet data and provides schema options similar to those for JSON and CSV files.

Security

Authentication

Teradata supports three authentication mechanisms in EVS on Teradata Cloud (VCE) and Teradata Factory (VMware and IntelliFlex) at the time of publishing this post: basic authentication (TD2), LDAP, and JWT token-based authentication.

Service URL

You need the Vector Store service URL to access the Vector Store APIs. The base_url in Teradata Cloud (VCE) is obtained from the provisioned site, and the endpoint information varies based on whether it is used as a stand-alone service or as part of AI Studio, as follows:

  • Stand-alone: https://[site-id].private.cloud.teradata.com
  • AI Studio: https://[site-id].private.cloud.teradata.com/one-td

For VantageCore Intelliflex or Vantage on VMware, base_url is the same as the AIWB login URL — for example, https://customer-aiwb-ui-login.com

For Artemis, base_url is the AI Studio deployment URL — for example, https://ai-studio-login.com/one-td

Access EVS from applications

You can access the EVS service from the service node using localhost and from within the system using the system IP address. However, to access the service from an external client, use the private link IP address.

Vector Store service listens on port 443, whereas the Teradata Database listens on port 1025. Both of these ports must be enabled through the private link.

Prerequisites

This section shows the high-level prerequisites to start using Teradata Enterprise Vector Store at the time of publishing this post. However, you should communicate Teradata you plan to use EVS in case there are additional requirements. You should also review the online documentation listed in the Documentation for Teradata Enterprise Vector Store section.

Client software

  • Python 3.9 or later. You also need these Python packages:
    • teradatagenai – Teradata Package for Generative AI.
    • teradataml – Teradata Package for Python.
  • The development environment is compatible with the Teradata Package for Generative AI: Jupyter and Spyder.
  • Minimum OS:
    • Windows 7 (64-bit) or later
    • macOS 10.9 (64-bit) or later
    • Red Hat 7 or later
    • Ubuntu 16.04 or later
    • CentOS 7 or later
    • SLES 12 or later

Database software

Database version 20.00.29.xx or above, and the Vector Store must be enabled.

Teradata Cloud

Teradata supports EVS as a stand-alone deployment or as part of an AI Studio deployment. However, for a given Teradata Cloud (VCE) instance, Teradata supports only one deployment choice.

Teradata Factory

Request Teradata for the prerequisites to use Enterprise Vector Store on-premises.

Documentation for Teradata Enterprise Vector Store

Teradata Enterprise Vector Store and its associated packages constantly evolve their capabilities. For the most up-to-date information on each component and its features, review the latest version of the Teradata online documentation. There you can find:

Additionally, Teradata maintains Python packages on PyPI, the official repository for Python packages. For example, you can find teradatagenai, teradataml, teradatamlwidgets and tdapiclient among those packages.

Separately, you can read the following posts to learn more about Teradata OTF capabilities:

This blog also includes several relevant posts on working with OTF, under the OTF category, independently of whether you use Teradata or another engine.


References

Teradata. (n.d.).


Comments

Leave a Reply

Your email address will not be published. Required fields are marked *