Skip to content

daft.functions.regexp_replace#

regexp_replace #

regexp_replace(expr: Expression, pattern: str | Expression, replacement: str | Expression) -> Expression

Replaces all occurrences of a regex pattern in a string column with a replacement string.

In the replacement, \1-\9/$1-$9 reference capture groups and \0/$0 the whole match; ${name}/$name reference a named group. Only one digit follows a backslash, so \10 is group 1 then a literal 0 and ${10} reaches group 10. \\ emits a literal backslash and $$ a literal $; any other backslash is literal, so \n is a backslash followed by n, not a newline.

Parameters:

Name Type Description Default
expr Expression

The string expression to be replaced

required
pattern str | Expression

The pattern to replace

required
replacement str | Expression

The replacement string

required

Returns:

Name Type Description
Expression Expression

a String expression with patterns replaced by the replacement string

Examples:

1
2
3
4
5
>>> import daft
>>> from daft.functions import regexp_replace
>>>
>>> df = daft.from_pydict({"data": ["foo", "fooo", "foooo"]})
>>> df.with_column("replace", regexp_replace(df["data"], r"o+", "a")).collect()
╭────────┬─────────╮
│ data   ┆ replace │
│ ---    ┆ ---     │
│ String ┆ String  │
╞════════╪═════════╡
│ foo    ┆ fa      │
├╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌┤
│ fooo   ┆ fa      │
├╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌┤
│ foooo  ┆ fa      │
╰────────┴─────────╯
(Showing first 3 of 3 rows)

Capture groups and a literal backslash:

1
2
3
>>> df = daft.from_pydict({"data": ["2024-01-31"]})
>>> df = df.select(regexp_replace(df["data"], r"(\d+)-(\d+)-(\d+)", r"\3\\\2\\\1"))
>>> df.to_pydict()["data"]
['31\\01\\2024']
Source code in daft/functions/str.py
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
def regexp_replace(
    expr: Expression,
    pattern: str | Expression,
    replacement: str | Expression,
) -> Expression:
    r"""Replaces all occurrences of a regex pattern in a string column with a replacement string.

    In the replacement, `\1`-`\9`/`$1`-`$9` reference capture groups and `\0`/`$0` the
    whole match; `${name}`/`$name` reference a named group. Only one digit follows a
    backslash, so `\10` is group 1 then a literal `0` and `${10}` reaches group 10.
    `\\` emits a literal backslash and `$$` a literal `$`; any other backslash is
    literal, so `\n` is a backslash followed by `n`, not a newline.

    Args:
        expr: The string expression to be replaced
        pattern: The pattern to replace
        replacement: The replacement string

    Returns:
        Expression: a String expression with patterns replaced by the replacement string

    Examples:
        >>> import daft
        >>> from daft.functions import regexp_replace
        >>>
        >>> df = daft.from_pydict({"data": ["foo", "fooo", "foooo"]})
        >>> df.with_column("replace", regexp_replace(df["data"], r"o+", "a")).collect()
        ╭────────┬─────────╮
        │ data   ┆ replace │
        │ ---    ┆ ---     │
        │ String ┆ String  │
        ╞════════╪═════════╡
        │ foo    ┆ fa      │
        ├╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌┤
        │ fooo   ┆ fa      │
        ├╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌┤
        │ foooo  ┆ fa      │
        ╰────────┴─────────╯
        <BLANKLINE>
        (Showing first 3 of 3 rows)

        Capture groups and a literal backslash:

        >>> df = daft.from_pydict({"data": ["2024-01-31"]})
        >>> df = df.select(regexp_replace(df["data"], r"(\d+)-(\d+)-(\d+)", r"\3\\\2\\\1"))
        >>> df.to_pydict()["data"]
        ['31\\01\\2024']

    """
    return Expression._call_builtin_scalar_fn("regexp_replace", expr, pattern, replacement)