bert-base-uncased
这是一款面向英文场景的不区分大小写的BERT基础预训练大模型底座,可微调适配文本分类、智能问答等各类文本处理下游任务。
这个项目值得继续研究吗?
这是一款面向英文场景的不区分大小写的BERT基础预训练大模型底座,可微调适配文本分类、智能问答等各类文本处理下游任务。
- 解决什么问题
- 企业落地英文文本分类、内容审核、智能问答、语义匹配等业务时,从零训练模型成本高、周期长、效果难保障,该模型可作为底座降低落地门槛。
- 适合什么团队
- 有英文文本类业务需求、具备基础AI研发能力的企业团队,可使用该模型快速搭建适配自身业务的文本处理能力。
- 使用前注意
- 该模型仅支持英文场景,训练数据带来的性别等偏见会传导至微调版本,不可直接用于文本生成场景,需微调后适配业务,Apache-2.0许可可商用。
本页用于缩短初步筛选时间,不构成技术、采购或法律结论。 正式使用前请在真实业务数据上验证,并以官方说明与许可证为准。
从官方资料看清能力、部署与采用边界
以下内容依据项目公开 README 或模型卡翻译整理,代码、命令和产品名保持原样。
项目定位
bert-base-uncased是谷歌2018年发布的BERT系列基础预训练模型,BERT是基于Transformer(当前大语言模型普遍采用的基础技术架构)的双向语义理解预训练模型系列,本版本为不区分英文大小写的基础款,属于大模型底座类产品,核心定位是为各类英文文本处理任务提供通用的语义理解能力支撑,无需企业从零开始训练文本理解模型。本模型卡片由Hugging Face团队维护,原生支持PyTorch、TensorFlow、JAX等主流AI开发框架调用。
核心能力
该模型基于英文书籍语料库、英文维基百科的海量公开无标注文本预训练完成,内置两类基础能力:一是掩码填空(输入带空缺的英文句子,模型预测空缺位置的合适内容),二是下句预测(判断两个英文句子是否为上下文连贯的关系)。 以上基础能力不直接适配业务场景,可通过微调(在预训练模型基础上用业务标注数据做小量训练)适配文本分类、实体识别、智能问答、语义相似度匹配等下游业务任务。该模型参数量为1.1亿,在通用英文语义理解测试集GLUE上的平均准确率为79.6%。 需要注意的是,该模型不支持文本生成类场景,且训练数据中隐含的性别等刻板印象偏见会传递到所有微调后的版本,比如测试中模型预测男性职业多为木匠、技工,女性职业多为护士、服务员,业务落地时需要做对应校验。
典型使用方式
针对普通业务场景,建议优先在Hugging Face模型库中查找对应任务已经微调完成的BERT衍生版本,无需自行训练即可直接使用。如果没有匹配的现成模型,再基于本模型做自定义微调。 技术团队可通过Hugging Face Transformers库快速调用该模型,几行代码即可实现两类基础能力的调用,也可直接调用模型输出任意英文文本的语义特征,用于训练业务所需的自定义分类器等模块。
部署与适配要求
该模型参数量较小,普通商用服务器甚至配置稍高的个人电脑即可运行,无需超高配置的专业GPU。同时支持CoreML、ONNX、Rust等多种格式导出,可适配云侧、端侧、边缘侧等不同部署环境。 如果需要做自定义微调,仅需准备对应业务场景的标注数据集,标注量根据任务复杂度从几百条到几千条不等,相比从零训练模型可节省大量标注和训练成本。
许可与采用建议
该模型采用Apache-2.0开源许可,可免费商用,没有开源传染风险,二次开发后的代码无需强制开源。 采用建议如下:1. 仅适用于英文文本类场景,中文场景请选择同系列的bert-base-chinese版本;2. 适合文本分类、内容审核、语义匹配、智能问答等理解类场景,文本生成类场景请选择GPT系列模型;3. 落地前需针对业务场景做偏见校验,避免输出不符合公序良俗的结果。
官方资料与来源
- transformers
- pytorch
- tf
- jax
- rust
- coreml
- onnx
- safetensors
- bert
- fill-mask
- exbert
- en

核对上游原始说明节选
任务类型:fill-mask
BERT base model (uncased)
Pretrained model on English language using a masked language modeling (MLM) objective. It was introduced in this paper and first released in this repository. This model is uncased: it does not make a difference between english and English.
Disclaimer: The team releasing BERT did not write a model card for this model so this model card has been written by the Hugging Face team.
Model description
BERT is a transformers model pretrained on a large corpus of English data in a self-supervised fashion. This means it was pretrained on the raw texts only, with no humans labeling them in any way (which is why it can use lots of publicly available data) with an automatic process to generate inputs and labels from those texts. More precisely, it was pretrained with two objectives:
the entire masked sentence through the model and has to predict the masked words. This is different from traditional recurrent neural networks (RNNs) that usually see the words one after the other, or from autoregressive models like GPT which internally masks the future tokens. It allows the model to learn a bidirectional representation of the sentence.
they correspond to sentences that were next to each other in the original text, sometimes not. The model then has to predict if the two sentences were following each other or not.
- Masked language modeling (MLM): taking a sentence, the model randomly masks 15% of the words in the input then run
- Next sentence prediction (NSP): the models concatenates two masked sentences as inputs during pretraining. Sometimes
This way, the model learns an inner representation of the English language that can then be used to extract features useful for downstream tasks: if you have a dataset of labeled sentences, for instance, you can train a standard classifier using the features produced by the BERT model as inputs.
Model variations
BERT has originally been released in base and large variations, for cased and uncased input text. The uncased models also strips out an accent markers. Chinese and multilingual uncased and cased versions followed shortly after. Modified preprocessing with whole word masking has replaced subpiece masking in a following work, with the release of two models. Other 24 smaller models are released afterward.
The detailed release history can be found on the google-research/bert readme on github.
| Model | #params | Language | |------------------------|--------------------------------|-------| | bert-base-uncased | 110M | English | | bert-large-uncased | 340M | English | sub | bert-base-cased | 110M | English | | bert-large-cased | 340M | English | | bert-base-chinese | 110M | Chinese | | bert-base-multilingual-cased | 110M | Multiple | | bert-large-uncased-whole-word-masking | 340M | English | | bert-large-cased-whole-word-masking | 340M | English |
Intended uses & limitations
You can use the raw model for either masked language modeling or next sentence prediction, but it's mostly intended to be fine-tuned on a downstream task. See the model hub to look for fine-tuned versions of a task that interests you.
Note that this model is primarily aimed at being fine-tuned on tasks that use the whole sentence (potentially masked) to make decisions, such as sequence classification, token classification or question answering. For tasks such as text generation you should look at model like GPT2.
How to use
You can use this model directly with a pipeline for masked language modeling:
>>> from transformers import pipeline
>>> unmasker = pipeline('fill-mask', model='bert-base-uncased')
>>> unmasker("Hello I'm a [MASK] model.")
[{'sequence': "[CLS] hello i'm a fashion model. [SEP]",
'score': 0.1073106899857521,
'token': 4827,
'token_str': 'fashion'},
{'sequence': "[CLS] hello i'm a role model. [SEP]",
'score': 0.08774490654468536,
'token': 2535,
'token_str': 'role'},
{'sequence': "[CLS] hello i'm a new model. [SEP]",
'score': 0.05338378623127937,
'token': 2047,
'token_str': 'new'},
{'sequence': "[CLS] hello i'm a super model. [SEP]",
'score': 0.04667217284440994,
'token': 3565,
'token_str': 'super'},
{'sequence': "[CLS] hello i'm a fine model. [SEP]",
'score': 0.027095865458250046,
'token': 2986,
'token_str': 'fine'}]Here is how to use this model to get the features of a given text in PyTorch:
from transformers import BertTokenizer, BertModel
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
model = BertModel.from_pretrained("bert-base-uncased")
text = "Replace me by any text you'd like."
encoded_input = tokenizer(text, return_tensors='pt')
output = model(**encoded_input)and in TensorFlow:
from transformers import BertTokenizer, TFBertModel
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
model = TFBertModel.from_pretrained("bert-base-uncased")
text = "Replace me by any text you'd like."
encoded_input = tokenizer(text, return_tensors='tf')
output = model(encoded_input)Limitations and bias
Even if the training data used for this model could be characterized as fairly neutral, this model can have biased predictions:
>>> from transformers import pipeline
>>> unmasker = pipeline('fill-mask', model='bert-base-uncased')
>>> unmasker("The man worked as a [MASK].")
[{'sequence': '[CLS] the man worked as a carpenter. [SEP]',
'score': 0.09747550636529922,
'token': 10533,
'token_str': 'carpenter'},
{'sequence': '[CLS] the man worked as a waiter. [SEP]',
'score': 0.0523831807076931,
'token': 15610,
'token_str': 'waiter'},
{'sequence': '[CLS] the man worked as a barber. [SEP]',
'score': 0.04962705448269844,
'token': 13362,
'token_str': 'barber'},
{'sequence': '[CLS] the man worked as a mechanic. [SEP]',
'score': 0.03788609802722931,
'token': 15893,
'token_str': 'mechanic'},
{'sequence': '[CLS] the man worked as a salesman. [SEP]',
'score': 0.037680890411138535,
'token': 18968,
'token_str': 'salesman'}]
>>> unmasker("The woman worked as a [MASK].")
[{'sequence': '[CLS] the woman worked as a nurse. [SEP]',
'score': 0.21981462836265564,
'token': 6821,
'token_str': 'nurse'},
{'sequence': '[CLS] the woman worked as a waitress. [SEP]',
'score': 0.1597415804862976,
'token': 13877,
'token_str': 'waitress'},
{'sequence': '[CLS] the woman worked as a maid. [SEP]',
'score': 0.1154729500412941,
'token': 10850,
'token_str': 'maid'},
{'sequence': '[CLS] the woman worked as a prostitute. [SEP]',
'score': 0.037968918681144714,
'token': 19215,
'token_str': 'prostitute'},
{'sequence': '[CLS] the woman worked as a cook. [SEP]',
'score': 0.03042375110089779,
'token': 5660,
'token_str': 'cook'}]This bias will also affect all fine-tuned versions of this model.
Training data
The BERT model was pretrained on BookCorpus, a dataset consisting of 11,038 unpublished books and English Wikipedia (excluding lists, tables and headers).
Training procedure
Preprocessing
The texts are lowercased and tokenized using WordPiece and a vocabulary size of 30,000. The inputs of the model are then of the form:
[CLS] Sentence A [SEP] Sentence B [SEP]With probability 0.5, sentence A and sentence B correspond to two consecutive sentences in the original corpus, and in the other cases, it's another random sentence in the corpus. Note that what is considered a sentence here is a consecutive span of text usually longer than a single sentence. The only constrain is that the result with the two "sentences" has a combined length of less than 512 tokens.
The details of the masking procedure for each sentence are the following:
- 15% of the tokens are masked.
- In 80% of the cases, the masked tokens are replaced by [MASK].
- In 10% of the cases, the masked tokens are replaced by a random token (different) from the one they replace.
- In the 10% remaining cases, the masked tokens are left as is.
Pretraining
The model was trained on 4 cloud TPUs in Pod configuration (16 TPU chips total) for one million steps with a batch size of 256. The sequence length was limited to 128 tokens for 90% of the steps and 512 for the remaining 10%. The optimizer used is Adam with a learning rate of 1e-4, \\(\beta{1} = 0.9\\) and \\(\beta{2} = 0.999\\), a weight decay of 0.01, learning rate warmup for 10,000 steps and linear decay of the learning rate after.
Evaluation results
When fine-tuned on downstream tasks, this model achieves the following results:
Glue test results:
| Task | MNLI-(m/mm) | QQP | QNLI | SST-2 | CoLA | STS-B | MRPC | RTE | Average | |:----:|:-----------:|:----:|:----:|:-----:|:----:|:-----:|:----:|:----:|:-------:| | | 84.6/83.4 | 71.2 | 90.5 | 93.5 | 52.1 | 85.8 | 88.9 | 66.4 | 79.6 |
BibTeX entry and citation info
@article{DBLP:journals/corr/abs-1810-04805,
author = {Jacob Devlin and
Ming{-}Wei Chang and
Kenton Lee and
Kristina Toutanova},
title = {{BERT:} Pre-training of Deep Bidirectional Transformers for Language
Understanding},
journal = {CoRR},
volume = {abs/1810.04805},
year = {2018},
url = {http://arxiv.org/abs/1810.04805},
archivePrefix = {arXiv},
eprint = {1810.04805},
timestamp = {Tue, 30 Oct 2018 20:39:56 +0100},
biburl = {https://dblp.org/rec/journals/corr/abs-1810-04805.bib},
bibsource = {dblp computer science bibliography, https://dblp.org}
}