Knowledgebase
The knowledgebase allows you to upload, process, and search through your documents with intelligent vector search and chunking. Documents of various types are automatically processed, embedded, and made searchable. Your documents are intelligently chunked, and you can view, edit, and search through them using natural language queries.
Upload and Processing
Simply upload your documents to get started. TradingGoose automatically processes them in the background, extracting text, creating embeddings, and breaking them into searchable chunks.
The system handles the entire processing pipeline for you:
- Text Extraction: Content is extracted from your documents using specialized parsers for each file type
- Intelligent Chunking: Documents are broken into meaningful chunks with configurable size and overlap
- Embedding Generation: Vector embeddings are created for semantic search capabilities
- Processing Status: Track the progress as your documents are processed
Supported File Types
TradingGoose supports PDF, Word (DOC/DOCX), plain text (TXT), Markdown (MD), HTML, Excel (XLS/XLSX), PowerPoint (PPT/PPTX), and CSV files. Files can be up to 100MB each, with optimal performance for files under 50MB. You can upload multiple documents simultaneously, and PDF files include OCR processing for scanned documents.
Viewing and Editing Chunks
Once your documents are processed, you can view and edit the individual chunks. This gives you full control over how your content is organized and searched. The chunks view displays each chunk's text content, its index within the document, and any associated tags.
Chunk Configuration
- Default chunk size: 1,024 estimated tokens for the text chunker, using approximately four characters per token
- Configurable range: 100-4,000 for the chunk-size setting; this is not a character limit
- Overlap: The text chunker prepends up to 200 words from the previous chunk by default; this can increase the final chunk size
- Hierarchical splitting: Respects document structure (sections, paragraphs, sentences)
These units describe the text chunker. Structured formats use format-specific chunkers, so the setting is not a universal bound on final text length.
Editing Capabilities
- Edit chunk content: Modify the text content of individual chunks
- Add chunks: Create additional chunks when needed; dedicated merge/split controls are not available
- Document tags: Organize documents using their tag fields
- Bulk operations: Enable, disable, or delete selected chunks
Advanced PDF Processing
For PDF documents, TradingGoose offers enhanced processing capabilities:
OCR Support
When configured with Azure or Mistral OCR:
- Scanned document processing: Extract text from image-based PDFs
- Mixed content handling: Process PDFs with both text and images
- High accuracy: Advanced AI models ensure accurate text extraction
Using The Knowledge Block in Workflows
Once your documents are processed, add a Knowledge block, select a knowledgebase, and configure a search query. To use the results for Retrieval-Augmented Generation (RAG), explicitly reference the block's results in the downstream Agent prompt, for example <knowledge.results> when the block is named knowledge. Connecting the blocks alone does not inject retrieved content.
The Knowledge block supports query-only, tag-only, and combined searches. In Advanced mode, configure text-equality tag filters; the block converts them to the API's filters contract. Repeated values for one tag use OR, while different tags use AND. Unknown tags and unsupported operators are rejected. Tags organize retrieval but are not an access-control boundary. See Tags and Filtering.
Knowledge Block Features
- Semantic search: Find relevant content using natural language queries
- Context integration: Explicitly reference retrieved chunks in agent prompts
- Dynamic retrieval: Search happens in real-time during workflow execution
- Relevance scoring: Results ranked by semantic similarity
Integration Options
- System prompts: Provide context to your AI agents
- Dynamic context: Search and include relevant information during conversations
- Multi-document search: Query across your entire knowledgebase
- Filtered search API: Combine vector search with document tags through the API's
filtersfield; the block limitation above still applies
Vector Search Technology
TradingGoose uses vector search powered by pgvector to understand the meaning and context of your content:
Semantic Understanding
- Contextual search: Finds relevant content even when exact keywords don't match
- Concept-based retrieval: Understands relationships between ideas
- Multi-language support: Works across different languages
- Synonym recognition: Finds related terms and concepts
Search Capabilities
- Natural language queries: Ask questions in plain English
- Similarity search: Find conceptually similar content
- Tag-filtered vector search: Restrict semantic search to tagged documents; this is not keyword/vector hybrid search
- Configurable results: Request 1-100 chunks with
topK(default: 10). Similarity thresholds are selected internally and are not a user setting
Document Management
Organization Features
- Bulk upload: Upload multiple files at once via the asynchronous API
- Processing status: Real-time updates on document processing
- Search and filter: Find documents quickly in large collections
- Metadata tracking: Automatic capture of file information and processing details
Security and Privacy
- Secure storage: Documents stored with enterprise-grade security
- Access control: Workspace-based permissions
- Processing isolation: Each workspace has isolated document processing
- Document removal: Delete documents when they are no longer needed; configurable automatic document-retention policies are not available
Getting Started
- Navigate to your knowledgebase: Access from your workspace sidebar
- Upload documents: Drag and drop or select files to upload
- Monitor processing: Watch as documents are processed and chunked
- Explore chunks: View and edit the processed content
- Add to workflows: Use the Knowledge block to integrate with your AI agents
The knowledgebase transforms your static documents into an intelligent, searchable resource that your AI workflows can leverage for more informed and contextual responses.