Model reference · open weights

SecureBERT2.0-biencoder

Available as managed deployment Embeddings cisco-ai Embeddings 1 variants 7k dl/mo

SecureBERT2.0-biencoder is an open-weight embedding model from cisco-ai. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released bycisco-ai
TypeEmbedding models
TaskEmbeddings
Parameters (lead)149M
Context8k tokens
Runs withsentence-transformers
Based oncisco-ai/SecureBERT2.0-base
Released2025-10-06
Popularity7k downloads / month
LicenceOpen weights

About

What SecureBERT2.0-biencoder is

The SecureBERT 2.0 Bi-Encoder is a cybersecurity-domain sentence-similarity and document-embedding model fine-tuned from SecureBERT 2.0. It independently encodes queries and documents into a shared vector space for semantic search, information retrieval, and cybersecurity knowledge retrieval.


Read the full model card

Model Details

Model Description

  • Developed by: Cisco AI
  • Model type: Bi-Encoder (Sentence Transformer)
  • Architecture: ModernBERT backbone with dual encoders
  • Max sequence length: 1024 tokens
  • Output dimension: 768
  • Language: English
  • License: Apache-2.0
  • Finetuned from: cisco-ai/SecureBERT2.0-base

Uses

Direct Use

  • Semantic search and document similarity in cybersecurity corpora
  • Information retrieval and ranking for threat intelligence reports, advisories, and vulnerability notes
  • Document embedding for retrieval-augmented generation (RAG) and clustering

Downstream Use

  • Threat intelligence knowledge graph construction
  • Cybersecurity QA and reasoning systems
  • Security operations center (SOC) data mining

Out-of-Scope Use

  • Non-technical or general-domain text similarity
  • Generative or conversational tasks

Model Architecture

The Bi-Encoder encodes queries and documents independently into a joint vector space. This architecture enables scalable approximate nearest-neighbor search for candidate retrieval and semantic ranking.


Datasets

Fine-Tuning Datasets

Dataset CategoryNumber of Records
Cybersecurity QA corpus43 000
Security governance QA corpus60 000
Cybersecurity instruction–response corpus25 000
Cybersecurity rules corpus (evaluation)5 000
Dataset Descriptions
  • Cybersecurity QA corpus: 43 k question–answer pairs, reports, and technical documents covering network security, malware analysis, cryptography, and cloud security.
  • Security governance QA corpus: 60 k expert-curated governance and compliance QA pairs emphasizing clear, validated responses.
  • Cybersecurity instruction–response corpus: 25 k instructional pairs enabling reasoning and instruction-following.
  • Cybersecurity rules corpus: 5 k structured policy and guideline records used for evaluation.

How to Get Started with the Model

Using Sentence Transformers

pip install -U sentence-transformers

Run Model to Encode

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("cisco-ai/SecureBERT2.0-biencoder")

sentences = [
    "How would you use Amcache analysis to detect fileless malware?",
    "Amcache analysis provides forensic artifacts for detecting fileless malware ...",
    "To capture and display network traffic"
]

embeddings = model.encode(sentences)
print(embeddings.shape)

Compute Similarity

from sentence_transformers import util
similarity = util.cos_sim(embeddings, embeddings)
print(similarity)


Framework Versions

  • python: 3.10.10
  • sentence_transformers: 5.0.0
  • transformers: 4.52.4
  • PyTorch: 2.7.0+cu128
  • accelerate: 1.9.0
  • datasets: 3.6.0

Training Details

Example Schema
FieldTypeDescription
sentence_0stringQuery or short text input
sentence_1stringCandidate or document text
labelfloatSimilarity score (1.0 = relevant)
Example Samples
sentence_0sentence_1label
Under what circumstances does attribution bias distort intrusion linking?Attribution bias in intrusion linking occurs when analysts allow preconceived notions, organizational pressures, or cognitive shortcuts to influence their assessment of attack origins and relationships between incidents...1.0
How can you identify store buffer bypass speculation artifacts?Store buffer bypass speculation artifacts represent side-channel vulnerabilities that exploit speculative execution to leak sensitive information...1.0

Training Objective and Loss

The model was optimized to maximize semantic similarity between relevant cybersecurity text pairs using contrastive learning.

Loss Parameters
{
    "scale": 20.0,
    "similarity_fct": "cos_sim"
}

Reference

@article{aghaei2025securebert,
  title={SecureBERT 2.0: Advanced Language Model for Cybersecurity Intelligence},
  author={Aghaei, Ehsan and Jain, Sarthak and Arun, Prashanth and Sambamoorthy, Arjun},
  journal={arXiv preprint arXiv:2510.00240},
  year={2025}
}

Model Card Authors

Cisco AI

Model Card Contact

For inquiries, please contact ai-threat-intel@cisco.com

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys securebert2-0-biencoder for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (securebert2-0-biencoder below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/embeddings \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"securebert2-0-biencoder","input":"text to embed"}'

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms