# Bulk Loader transaction scope (blank node identification)

**URL:** <https://discuss.dgraph.io/t/bulk-loader-transaction-scope-blank-node-identification/12548>\
**Category:** Dgraph\
**Tags:** kind:question, dgraph, kind:bug\
**Created:** [January 29, 2021, 4:57pm UTC](https://discuss.dgraph.io/t/bulk-loader-transaction-scope-blank-node-identification/12548 "2021-01-29T16:57:08Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![apete](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/apete/32/5696_2.png) [@apete](https://discuss.dgraph.io/u/apete)\
**Post date:** [January 29, 2021, 4:57pm UTC](https://discuss.dgraph.io/t/bulk-loader-transaction-scope-blank-node-identification/12548/1 "2021-01-29T16:57:08Z")

</div>

When programmatically inserting mutations there is the concept of a transaction and the blank nodes from the mutations are recognised within the transaction. What does this transalte to when using the Bulk Loader? Is there something corresponding to a transaction? Are blank nodes in any way recognised between triples?

I am aware of the `--xidmap` option, but have been unable to use it. It consumes way too much memory. As I understand it, it is a disk based cache and should be able to have limited RAM usage. Whenever I run the bulk loader with the `--xidmap` option it consumes all available memory and eventually crashes.

This post

> [@Bulk loader xidmap memory optimization](http://discuss.hypermode.com/t/bulk-loader-xidmap-memory-optimization/7496):
>
> While bulk loading a dataset with a high number of unique xids, the bulk loader goes out of memory in the mapper phase. We keep a sharded map (XidMap) of xids to their corresponding id so that multiple instances of the same blank nodes get the same id. The collected memory required by the shared map and all the xids increases significantly. We propose that we introduce a command-line option to the bulk loader, --limitMemory. This option would allow the user to limit the memory required by the x…

discuss a possible option `--limitMemory`. There is no such option in the current version, right?

With or without such an option there seems to be a problem with offloading cached id:s to disk so that memory can be limited. Is there perhaps a known bug here?

That `--xidmap` option does precisely what we need, but we can’t use it.

---

<div class="post-metadata">

**Author:** ![MichelDiz](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/micheldiz/32/11873_2.png) [@MichelDiz](https://discuss.dgraph.io/u/MichelDiz)\
**Post date:** [January 29, 2021, 5:33pm UTC](https://discuss.dgraph.io/t/bulk-loader-transaction-scope-blank-node-identification/12548/2 "2021-01-29T17:33:15Z")

</div>

> [@apete](#):
>
> What does this transalte to when using the Bulk Loader?

There’s no transaction in the Bulk Loader. It is used only once to populate the cluster.

> [@apete](#):
>
> Are blank nodes in any way recognised between triples?

A blank node is just an identifier. But you can store that in the node itself by using `--store_xids` flag.  
That can be used in upsert queries(block) and also in Liveload.

```auto
➜ ~ dgraph bulk -h | grep xid
      --store_xids Generate an xid edge for each node.
      --xidmap string Directory to store xid to uid mapping

➜ ~ dgraph live -h | grep xid
  -U, --upsertPredicate string run in upsertPredicate mode. the value would be used to store blank nodes as an xid
  -x, --xidmap string Directory to store xid to uid mapping

```

> [@apete](#):
>
> `--xidmap` option it consumes all available memory and eventually crashes.

Try to use `--store_xids`.

---

<div class="post-metadata">

**Author:** ![apete](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/apete/32/5696_2.png) [@apete](https://discuss.dgraph.io/u/apete)\
**Post date:** [January 29, 2021, 5:58pm UTC](https://discuss.dgraph.io/t/bulk-loader-transaction-scope-blank-node-identification/12548/3 "2021-01-29T17:58:53Z")

</div>

Are you suggesting that `--xidmap` works differently when combined with `--store_xids`? or are you recommending to use `--store_xids` instead of `--xidmap` and then do `upsert`:s?

Should interpret you answer as the `--xidmap` option won’t work with larger data sets?

---

<div class="post-metadata">

**Author:** ![MichelDiz](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/micheldiz/32/11873_2.png) [@MichelDiz](https://discuss.dgraph.io/u/MichelDiz)\
**Post date:** [January 29, 2021, 8:08pm UTC](https://discuss.dgraph.io/t/bulk-loader-transaction-scope-blank-node-identification/12548/4 "2021-01-29T20:08:56Z")

</div>

> [@apete](#):
>
> or are you recommending to use `--store_xids` instead of `--xidmap` and then do `upsert`:s?

This, yes.

> [@apete](#):
>
> Should interpret you answer as the `--xidmap` option won’t work with larger data sets?

No, but it depends. OOMs are normal, it happens when you don’t have the idea of data x resources. When you don’t know how much resource you need for that particular dataset.

For example, the data in the blog post [Loading close to 1M edges/sec into Dgraph - Dgraph Blog](https://dgraph.io/blog/post/bulkloader/) was around 150GB. I don’t remember exactly the size, but it was around that. With that in mind, a dataset of 150GB should not go OOM easily. With the configuration mentioned in the blog post.

But sure, some limitations(e.g limit memory usage, add a cool down and so on) could be good to avoid it. But the time to process would increase.

---

<div class="post-metadata">

**Author:** ![apete](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/apete/32/5696_2.png) [@apete](https://discuss.dgraph.io/u/apete)\
**Post date:** [January 31, 2021, 11:10am UTC](https://discuss.dgraph.io/t/bulk-loader-transaction-scope-blank-node-identification/12548/5 "2021-01-31T11:10:38Z")

</div>

We have 400G RAM available. When using the bulk loader without the `--xidmap` option it is well behaved and only uses a fraction of that. With the `--xidmap` option memory consumption never stops growing, and the process eventually crashes. That seems like faulty behaviour to me.

How much memory does the `--xidmap` option need to function properly? Is there some metric bytes per unique blank node?

---

<div class="post-metadata">

**Author:** ![MichelDiz](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/micheldiz/32/11873_2.png) [@MichelDiz](https://discuss.dgraph.io/u/MichelDiz)\
**Post date:** [January 31, 2021, 4:43pm UTC](https://discuss.dgraph.io/t/bulk-loader-transaction-scope-blank-node-identification/12548/6 "2021-01-31T16:43:50Z")

</div>

> [@apete](#):
>
> How much memory does the `--xidmap` option need to function properly? Is there some metric bytes per unique blank node?

Not sure

> [@apete](#):
>
> With the `--xidmap` option memory consumption never stops growing, and the process eventually crashes.

let me ping @Anurag and @ibrahim - Maybe it needs a refactoring to use jemalloc or something.

---

<div class="post-metadata">

**Author:** ![apete](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/apete/32/5696_2.png) [@apete](https://discuss.dgraph.io/u/apete)\
**Post date:** [February 1, 2021, 9:20am UTC](https://discuss.dgraph.io/t/bulk-loader-transaction-scope-blank-node-identification/12548/7 "2021-02-01T09:20:37Z")

</div>

Just updated to the very latest version and tried this again.

With the `--xidmap` option the MAP phase is half speed compared to not using it. It quickly claims 100G and within 40min it consumed the entire 400G and crashed.

Without the `--xidmap` option it initially claims 50G, and it grows much slower. Right now I can’t tell you the max memory level it reaches – but it doesn’t crash.
